The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman
The MAD Podcast with Matt Turck
11 HOURS AGO
The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman
The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman

The MAD Podcast with Matt Turck
11 HOURS AGO
Shownote
Shownote
AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act. In this episode of the MAD Podcast, Matt Turck sits down wi...
Highlights
Highlights
In this episode of the MAD Podcast, Matt Turck and Andrew Feldman, CEO of Cerebras, explore the critical shift in AI from training to inference, arguing that speed is the new bottleneck. Feldman explains why tokens per second per user is the key metric and how Cerebras' wafer-scale chip, with its massive SRAM, overcomes the memory limitations of traditional GPUs for faster, more responsive AI.
Chapters
Chapters
Cold open & Intro
00:00Why speed became the AI bottleneck
01:31Tokens per second per user, explained
02:32AI’s broadband moment and the Netflix analogy
03:16The AI chip landscape: GPUs, TPUs, Trainium, ASICs
04:35What is an ASIC?
06:36Nvidia, Groq, and the fast inference war
08:08OpenAI, Broadcom, and specialized silicon
09:16China, power, and sovereign AI infrastructure
12:10Is the AI infrastructure boom a bubble?
15:05The hidden bottlenecks: HBM, CoWoS, and 3nm
18:56Why agents are creating CPU demand
22:57Andrew Feldman’s path from SeaMicro to Cerebras
25:36Why Cerebras bet on AI in 2016
26:13SRAM vs. HBM: why inference is a memory problem
31:14What wafer-scale computing actually means
33:19The deep-tech “Everest” problem
34:28The moment the first Cerebras system worked
36:07Ringing the bell and surviving deep tech
36:49How a giant chip handles failure
39:08Why GPUs struggle with decode
41:22Prefill vs. decode explained
42:17The “100 HD movies” problem in AI inference
44:01How fast inference changes RL and training
45:04Reasoning models and why they cost more compute
48:08Verification, guardrails, and small models checking big models
50:08Multimodal AI and the path to video
52:37Cerebras’ business model: hardware, cloud, and API
53:51OpenAI’s 750MW inference deal
55:14Why data centers are measured in megawatts
55:36AWS Trainium + Cerebras decode
58:01Fast tokens as a cloud product
59:29Is CUDA still a moat?
1:00:52How TSMC helped Cerebras build the giant chip
1:03:53Why nobody cared in 2020
1:07:41Why chip supply chains are hard to diversify
1:08:15Why today’s AI models will be the worst you ever use
1:09:54What fast AI could do to SaaS
1:10:38Transcript
Transcript
Andrew Feldman: This is the largest chip built in the history of the computer industry. It's 58 times larger than a GPU. And for AI, bigger chips process information more quickly, and therefore you get answers in less time. For AI worm, big chips are undou...