Google's AI Infrastructure Chief, Amin Vahdat, on the Physics & Economics of Frontier AI

At 100,000-accelerator scale, something fails multiple times an hour, which is why Google's Chief Technologist for AI Infrastructure Amin Vahdat thinks FLOPS is a vanity metric. The metric that matters is what he calls goodput: the useful work a workload actually delivers through real failures. Amin walks through the calculus that split the TPU line into 8i and 8t for the first time, why the TPU's core primitives haven't changed since v1, and how Google and DeepMind co-design in the same rooms, intercepting chip architectures mid-flight before tape-out. He explains why long-horizon agents are sending demand for CPUs and storage through the roof alongside accelerators, how optical circuit switches reroute light to a spare rack in milliseconds, and why Google would rather wait on a utility than build its own gigawatt. We also cover orbital data centers and the multi-megawatt rack of 2036.

Hosted by Sonya Huang, Sequoia Capital

0:00 – Introduction

1:47 – What makes a data center an AI data center

5:30 – Goodput, not FLOPS: holding yourself accountable when something fails every hour

11:52 – Doubling token capacity every six months, and where the gains actually come from

15:32 – The TPU bet: from a contrarian call in 2013 to splitting 8i and 8t

23:30 – The case for and against co-design

26:11 – Shoulder to shoulder with DeepMind: intercepting chips mid-flight

34:16 – Long-horizon agents change the shape of the data center

37:50 – Optical circuit switching and the state of networking

42:35 – Power is the binding constraint: utilities, gigawatts, and sizing a data center

49:23 – Training vs. serving clusters, seven-year-old TPUs, and open standards

58:29 – Orbital data centers and the supercomputer of 2036

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论