Google's AI Infrastructure Chief, Amin Vahdat, on the Physics & Economics of Frontier AI
At 100,000-accelerator scale, something fails multiple times an hour, which is why Google's Chief Technologist for AI Infrastructure Amin Vahdat thinks FLOPS is a vanity metric. The metric that matters is what he calls goodput: the useful work a workload actually delivers through real failures. Amin walks through the calculus that split the TPU line into 8i and 8t for the first time, why the TPU's core primitives haven't changed since v1, and how Google and DeepMind co-design in the same rooms, intercepting chip architectures mid-flight before tape-out. He explains why long-horizon agents are sending demand for CPUs and storage through the roof alongside accelerators, how optical circuit switches reroute light to a spare rack in milliseconds, and why Google would rather wait on a utility than build its own gigawatt. We also cover orbital data centers and the multi-megawatt rack of 2036.
Hosted by Sonya Huang, Sequoia Capital
0:00 – Introduction
1:47 – What makes a data center an AI data center
5:30 – Goodput, not FLOPS: holding yourself accountable when something fails every hour
11:52 – Doubling token capacity every six months, and where the gains actually come from
15:32 – The TPU bet: from a contrarian call in 2013 to splitting 8i and 8t
23:30 – The case for and against co-design
26:11 – Shoulder to shoulder with DeepMind: intercepting chips mid-flight
34:16 – Long-horizon agents change the shape of the data center
37:50 – Optical circuit switching and the state of networking
42:35 – Power is the binding constraint: utilities, gigawatts, and sizing a data center
49:23 – Training vs. serving clusters, seven-year-old TPUs, and open standards
58:29 – Orbital data centers and the supercomputer of 2036