scripod.com

The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman

Shownote

AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act. In this episode of the MAD Podcast, Matt Turck sits down wi...

Highlights

In this episode of the MAD Podcast, Matt Turck and Andrew Feldman, CEO of Cerebras, explore the critical shift in AI from training to inference, arguing that speed is the new bottleneck. Feldman explains why tokens per second per user is the key metric and how Cerebras' wafer-scale chip, with its massive SRAM, overcomes the memory limitations of traditional GPUs for faster, more responsive AI.
00:00
Our chip is 58 times larger than a GPU.
01:36
Fast tokens become more productive and valuable
02:33
The right metric for AI inference speed is tokens per second per user.
03:20
Slow AI has no market
04:36
Building chips from scratch optimized for AI
06:36
ASICs are specialized chips optimized for specific tasks.
08:18
GPUs cannot handle fast inference
09:16
A multi-silicon ecosystem is healthy
12:13
China invests in power and open-source despite chip disadvantages
15:09
Ignore short-term fluctuations, focus on building.
20:57
Cerebras chips are 58x larger than GPUs with 3000x more memory bandwidth
25:36
Perfect timing often follows a decade of bad timing
29:27
We chose to build a radically better chip.
31:22
AI inference is memory-bound
33:19
We spent 18 months spending $8 million a month and couldn't make one.
34:28
Packaging was the real challenge
36:08
Team stared in disbelief at the running server
36:49
IPO marks a plateau for future success
39:16
Redundancy and cooling are key to reliability.
41:31
Cerebras chips are faster than GPUs for inference.
42:17
Inference has two steps: prefill and decode.
44:01
SRAM moves weights 2,500 times faster than Nvidia GPU
45:14
RL uses inference within training.
48:14
Reasoning in AI is like writing multiple drafts, requiring more compute.
50:11
Verification and guardrails require extra compute time
52:44
Cerebras is fastest on a Google multimodal model.
54:04
OpenAI deal shifts mix toward cloud
55:14
Cerebras provides up to 750 megawatts of power
55:45
Bottlenecks are CoWoS, data centers, and power.
58:04
Flexibility is key, so we offer both.
59:35
Test models on demand, then deploy long-term.
1:00:52
Nvidia's CUDA moat is shrinking
1:03:53
Solving a historic problem with no initial market interest
1:08:15
AI was a hobby, not production-critical
1:08:23
Chip design is tied to a specific fab's rules
1:09:54
Current models will soon seem primitive
1:10:44
AI can instantly build tools like Salesforce

Chapters

Cold open & Intro
00:00
Why speed became the AI bottleneck
01:31
Tokens per second per user, explained
02:32
AI’s broadband moment and the Netflix analogy
03:16
The AI chip landscape: GPUs, TPUs, Trainium, ASICs
04:35
What is an ASIC?
06:36
Nvidia, Groq, and the fast inference war
08:08
OpenAI, Broadcom, and specialized silicon
09:16
China, power, and sovereign AI infrastructure
12:10
Is the AI infrastructure boom a bubble?
15:05
The hidden bottlenecks: HBM, CoWoS, and 3nm
18:56
Why agents are creating CPU demand
22:57
Andrew Feldman’s path from SeaMicro to Cerebras
25:36
Why Cerebras bet on AI in 2016
26:13
SRAM vs. HBM: why inference is a memory problem
31:14
What wafer-scale computing actually means
33:19
The deep-tech “Everest” problem
34:28
The moment the first Cerebras system worked
36:07
Ringing the bell and surviving deep tech
36:49
How a giant chip handles failure
39:08
Why GPUs struggle with decode
41:22
Prefill vs. decode explained
42:17
The “100 HD movies” problem in AI inference
44:01
How fast inference changes RL and training
45:04
Reasoning models and why they cost more compute
48:08
Verification, guardrails, and small models checking big models
50:08
Multimodal AI and the path to video
52:37
Cerebras’ business model: hardware, cloud, and API
53:51
OpenAI’s 750MW inference deal
55:14
Why data centers are measured in megawatts
55:36
AWS Trainium + Cerebras decode
58:01
Fast tokens as a cloud product
59:29
Is CUDA still a moat?
1:00:52
How TSMC helped Cerebras build the giant chip
1:03:53
Why nobody cared in 2020
1:07:41
Why chip supply chains are hard to diversify
1:08:15
Why today’s AI models will be the worst you ever use
1:09:54
What fast AI could do to SaaS
1:10:38

Transcript

Andrew Feldman: This is the largest chip built in the history of the computer industry. It's 58 times larger than a GPU. And for AI, bigger chips process information more quickly, and therefore you get answers in less time. For AI worm, big chips are undou...