AI Infra Summit Day 2: Threads + Highlights
Today is a guest post from Dustin Sklavos, former AnandTech Senior Editor, product manager at Corsair, and technical writer at Tenstorrent. Dustin brings almost two decades of experience writing and working in the industry, managing hardware line revenues north of $100m, and routinely acts as a robust sail in the face of superficial marketing. While looking for an interesting yet new challenge in the industry, I’m glad to have him writing for More Than Moore this week at the AI Infra Summit.
For day 2, we spent the time prowling the show floor for some interesting and innovative hardware approaches, as well as some of the trends.
Server Cooling
I’m an absolute sucker for cooling solutions. I worked on the liquid cooling loop built into the second generation Corsair One systems while at Corsair (the prebuilt consumer high-end PC-in-a-small-form-factor product line) and spent time managing their liquid coolers. Also at AI hardware startup Tenstorrent I helped develop the first generation TT-QuietBox, their consumer level dual-PCIe card system solution. Even my home system has a (superfluous) custom liquid cooling loop. My last article as a journalist was about my first experience building a loop.
Liquid cooling in the data center is becoming essential - as racks of hardware go beyond the 12-25 kilowatts of the enterprise market, especially for the multi-kilowatt ASICs in the pipeline, there were water blocks all the way down. Especially on photonics too, as that will be absorbed by the leading ASIC players first.
There were two water block vendors at the show that each had their own sauce.
Malico was using fairly conventional copper waterblocks, but they were stitching the loop together with full copper heatpipes. There’s a tradeoff here: serviceability is obviously more difficult because you now have an entire assembly to deal with in one unit, and you also lose any of the permeability of conventional tubing, but it can improve thermal performance. The tradeoff between serviceability and performance is going to be one that data center solutions managers are going to be fighting through the next decade (just don’t ask about space data centers).
Phononics (not pictured) was a company on display that developed a sensor-based thermoelectric cooling (TEC) solutions able to effectively ‘lock’ the operating temperature of HBM stacks. The operating theory here is that when the thermal tolerance and temperature ranges for HBM have to be designed into more traditional air and liquid cooling solutions, ranges it can lead to wasting power on excessive cooling when it’s not essential. Phononics is able to use a series of sensors to adjust the performance of their TEC plate within milliseconds, maintaining a constant, stable operating temperature on the memory. It’s fascinating stuff, but in my experience, a TEC can introduce its own drawbacks (power consumption ironically being a big one), and that’s made it traditionally a nonstarter in the consumer space. They’re still fairly popular for certain industrial and embedded installations however. The data center? I’ll believe it when it’s at scale perhaps.
I also visited with Alloy Enterprises, recently acquired by Johnson Controls, and their cold plate technology has its own twist. They use laser-cut, bonded layers of copper or aluminum (never mixed, one or the other) to create their cold plates. These blocks are actually optimized for the hot spots on the processors they’re designed for, improving overall cooling performance. Their competitor, Corintis (not present at AI Infra Summit), operates under a similar principle of optimizing the cooling pathways and design for chip hot spots, but uses a production process analogous to 3D printing copper.
One of the last ones I saw was also one of the most eyebrow-raising. Molten Dynamics (not pictured) is a very early stage startup that has claimed to figure out how to use liquid metal instead of standard glycol-and-water coolant inside a cooling loop.
My using an alloy of gallium (probably an indium/tin mix), the metal acts as a better conductor than any water or oil based liquid in the loop. Instead of a physical pump, the solution employs a Lorentz force magnetic pump with no moving parts to move the coolant - this means passing a current through the metal while also externally applying a magnetic field, forcing the metal through the pipes. One big downside is that the thermal capacity of the metal alloy is less than 10% of water, meaning it can’t absorb as much thermal energy, and a faster pump rate might be need. Also leaking gallium is a hell of a lot riskier than leaking standard coolant. Molten Dynamics however is at least a novel solution geared toward cooling processors with TDPs well above a kilowatt.
Edge AI Devices
The run on Apple Mac Minis and Mac Studios for AI models has demonstrated at least a niche market for running those models locally on your home network. This is a market now tackled by x86 and Arm solutions, such as the AMD Strix Halo and the Nvidia DGX Spark. If you blink, the price might have increased. But the premise of having an AI model local in your device rather than the cloud is still a priority in the edge and embedded space. We saw a few such solutions.
Innodisk was showing off a variety of systems built for running AI locally, including an industrial solution meant for environments that require high thermal tolerances. Innodisk’s hardware employs their own networking solutions, allowing users to bridge them together, but the actual AI computation is delightfully agnostic. Innodisk is offering products with a range of processors, with some focus on the software layer to improve ease of use.
On the opposite end, the box from Tiiny was of particular note. Small enough to be portable, not quite low power enough to just run off of a host USB-C port but still a good start. For 30 watts and featuring 80GB of local memory (split between 32GB for the host SoC and 48GB for their dedicated AI processor), it allows for running models with upwards of 100B parameters.
What Tiiny got particularly right with their product was the UX. Their software has a clean, user-friendly interface, andruns on both Windows and macOS. It has an App Store-like component built into it that allows you to download models on demand, and then runs the models inside the software. What I saw was shockingly easy to use and other companies looking at developing agentic AI devices should take notice of what they’ve accomplished here. Hopefully we’ll be able to get one in for independent review. That being said, the people we spoke to at the booth said that they’re sold out for the next year, and now going after bigger business customers rather than the consumer market.
A Brief Word About RISC-V
Depending on who you ask, RISC-V seems to still be struggling to break through into the data center AI space. At first glance, the standard’s presence at AI Infra Summit was underwhelming. Depending on who you talk to, there’s either a large market and a lot of potential for it (especially as Arm continues to escalate hostilities against their own customers), or the technology just isn’t ever going to get there.
There was a corner of the show floor with a few tables for companies like SiFive and Akeana; Epic Semi had a proper booth and showed off a server solution that was RISC-V from top to bottom: the host CPU was based on RISC-V cores, and the AI processor was based on RISC-V cores.
What bears mentioning is that RISC-V technology was way more prevalent at the show than the corner of the AI Infra Summit floor made it appear. RISC-V cores in all shapes and sizes were integrated into various solutions. Remember that Tenstorrent uses RISC-V cores both in their AI architecture as well as their CPU offerings and Axelera is using RISC-V cores to control dataflow in their AI cards. And even Xcena’s MX1 CXL modules, built to expand memory and flash capacity, use over a thousand RISC-V cores to intelligently move and cache data.
About That CXL
With AI’s near bottomless thirst for storage and memory, CXL was also alive and kicking hard at AI Infra Summit. Marvell featured expansion solutions using CXL (the Structera family), and Xcena was showing off their MX1 modules and switches. Xcena presented their MX1 architecture at Hot Chips earlier this year. It turns out that with the reuse of hardware, either extended memory or a composable system is coming back into vogue.
Thanks for reading More Than Moore! This post is public so feel free to share it.
Conclusion
Probably one of the biggest takeaways I had from the AI Infra Summit show floor was just how diverse the technologies on tap are. AI is still a very open field, and as Professor Dave Patterson mentioned during the keynote, the processor architectures needed to support it are having to constantly evolve. That includes every piece of surrounding infrastructure.
It’s wild to watch all of this evolve in real time, seeing firms aggressively developing new solutions to feed this endlessly hungry beast, driving faster compute, faster and more reliable interconnects, higher density and faster memory and storage solutions, making it fun size for home use, optimizing power and serviceability and security, and on and on. I can’t imagine what opinion I’m going to have this time next year.