The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman
The MAD Podcast with Matt Turck
13 HOURS AGO
The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman
The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman

The MAD Podcast with Matt Turck
13 HOURS AGO
In this episode of the MAD Podcast, Matt Turck and Andrew Feldman, CEO of Cerebras, explore the critical shift in AI from training to inference, arguing that speed is the new bottleneck. Feldman explains why tokens per second per user is the key metric and how Cerebras' wafer-scale chip, with its massive SRAM, overcomes the memory limitations of traditional GPUs for faster, more responsive AI.
Feldman defines the AI chip landscape, contrasting general-purpose GPUs with specialized ASICs like Cerebras' wafer-scale chip, which is 58 times larger than a GPU. He explains that AI inference is memory-bound, and Cerebras uses on-chip SRAM to move model weights 2,500 times faster than HBM-based GPUs, solving the '100 HD movies' problem. This speed is crucial for reasoning models, agents, and reinforcement learning, where latency is amplified. Feldman argues that Nvidia's CUDA moat is shrinking for inference, as switching is easy. He discusses Cerebras' partnership with AWS, where Trainium handles prefill and Cerebras handles decode, and its 750MW deal with OpenAI. He concludes that fast AI will reshape SaaS by enabling instant tool creation and breaking down organizational silos, and that today's models will be the worst we ever use.
00:00
00:00
Our chip is 58 times larger than a GPU.
01:36
01:36
Fast tokens become more productive and valuable
02:33
02:33
The right metric for AI inference speed is tokens per second per user.
03:20
03:20
Slow AI has no market
04:36
04:36
Building chips from scratch optimized for AI
06:36
06:36
ASICs are specialized chips optimized for specific tasks.
08:18
08:18
GPUs cannot handle fast inference
09:16
09:16
A multi-silicon ecosystem is healthy
12:13
12:13
China invests in power and open-source despite chip disadvantages
15:09
15:09
Ignore short-term fluctuations, focus on building.
20:57
20:57
Cerebras chips are 58x larger than GPUs with 3000x more memory bandwidth
25:36
25:36
Perfect timing often follows a decade of bad timing
29:27
29:27
We chose to build a radically better chip.
31:22
31:22
AI inference is memory-bound
33:19
33:19
We spent 18 months spending $8 million a month and couldn't make one.
34:28
34:28
Packaging was the real challenge
36:08
36:08
Team stared in disbelief at the running server
36:49
36:49
IPO marks a plateau for future success
39:16
39:16
Redundancy and cooling are key to reliability.
41:31
41:31
Cerebras chips are faster than GPUs for inference.
42:17
42:17
Inference has two steps: prefill and decode.
44:01
44:01
SRAM moves weights 2,500 times faster than Nvidia GPU
45:14
45:14
RL uses inference within training.
48:14
48:14
Reasoning in AI is like writing multiple drafts, requiring more compute.
50:11
50:11
Verification and guardrails require extra compute time
52:44
52:44
Cerebras is fastest on a Google multimodal model.
54:04
54:04
OpenAI deal shifts mix toward cloud
55:14
55:14
Cerebras provides up to 750 megawatts of power
55:45
55:45
Bottlenecks are CoWoS, data centers, and power.
58:04
58:04
Flexibility is key, so we offer both.
59:35
59:35
Test models on demand, then deploy long-term.
1:00:52
1:00:52
Nvidia's CUDA moat is shrinking
1:03:53
1:03:53
Solving a historic problem with no initial market interest
1:08:15
1:08:15
AI was a hobby, not production-critical
1:08:23
1:08:23
Chip design is tied to a specific fab's rules
1:09:54
1:09:54
Current models will soon seem primitive
1:10:44
1:10:44
AI can instantly build tools like Salesforce