scripod.com

The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman

AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act. In this episode of the MAD Podcast, Matt Turck sits down with Andrew Feldman, co-founder and CEO of Cerebras, to explain why fast inference may define the next era of AI.
Cerebras is known for building a chip the size of a silicon wafer. But this conversation is not just about one company or one chip. It is a deep dive into the AI infrastructure stack: GPUs, ASICs, memory, HBM, SRAM, data centers, power, TSMC, AWS, OpenAI, agents, reasoning models, and why speed changes what AI products can become. Andrew explains why “tokens per second per user” matters, why generating a single word can require moving the equivalent of 100 HD movies through memory, why agents amplify latency, why GPUs struggle with certain inference workloads, and why fast AI may eventually reshape SaaS itself.
This is a reference conversation on fast inference, AI chips, and the next compute bottleneck.
(00:00) Cold open & Intro
(01:31) Why speed became the AI bottleneck
(02:32) Tokens per second per user, explained
(03:16) AI’s broadband moment and the Netflix analogy
(04:35) The AI chip landscape: GPUs, TPUs, Trainium, ASICs
(06:36) What is an ASIC?
(08:08) Nvidia, Groq, and the fast inference war
(09:16) OpenAI, Broadcom, and specialized silicon
(12:10) China, power, and sovereign AI infrastructure
(15:05) Is the AI infrastructure boom a bubble?
(18:56) The hidden bottlenecks: HBM, CoWoS, and 3nm
(22:57) Why agents are creating CPU demand
(25:36) Andrew Feldman’s path from SeaMicro to Cerebras
(26:13) Why Cerebras bet on AI in 2016
(31:14) SRAM vs. HBM: why inference is a memory problem
(33:19) What wafer-scale computing actually means
(34:28) The deep-tech “Everest” problem
(36:07) The moment the first Cerebras system worked
(36:49) Ringing the bell and surviving deep tech
(39:08) How a giant chip handles failure
(41:22) Why GPUs struggle with decode
(42:17) Prefill vs. decode explained
(44:01) The “100 HD movies” problem in AI inference
(45:04) How fast inference changes RL and training
(48:08) Reasoning models and why they cost more compute
(50:08) Verification, guardrails, and small models checking big models
(52:37) Multimodal AI and the path to video
(53:51) Cerebras’ business model: hardware, cloud, and API
(55:14) OpenAI’s 750MW inference deal
(55:36) Why data centers are measured in megawatts
(58:01) AWS Trainium + Cerebras decode
(59:29) Fast tokens as a cloud product
(01:00:52) Is CUDA still a moat?
(01:03:53) How TSMC helped Cerebras build the giant chip
(01:07:41) Why nobody cared in 2020
(01:08:15) Why chip supply chains are hard to diversify
(01:09:54) Why today’s AI models will be the worst you ever use
(01:10:38) What fast AI could do to SaaS