scripod.com

Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud

Shownote

Baseten CEO and co-founder Tuhin Srivastava sits down with Sarah Guo and Elad Gil to discuss the rapid growth of AI inference demand, Baseten’s 30x growth, and why inference is becoming the strategic “last market.” Tuhin Srivastava argues the application l...

Highlights

In this podcast, Baseten CEO Tuhin Srivastava discusses the company's explosive 30x growth in AI inference, driven by widespread adoption and the strategic importance of the application layer. He argues that companies with unique user data can build lasting value through specialized workflows and post-trained models, using examples like Abridge. The conversation also covers GPU supply constraints, the rise of multi-cloud infrastructure, and the evolving dynamics of long-term hardware contracts.
00:05
Baseten's 30x growth in AI inference
02:04
Application layer will persist because companies with unique user signals can encode value into workflows
05:57
Serving AI-native companies prepares us for enterprise needs
07:55
Customers prioritize capability over cost.
12:14
Running DeepSeek costs about 20% of Anthropic models in production.
13:07
Over 95% of tokens served on Baseten come from custom models.
14:22
Inference and post-training are increasingly linked
17:10
Prove product-market fit before custom models
18:35
Very little slack compute and high utilization
22:28
GPU contracts require 3-5 year commitments with prepayment
24:10
Software stickiness drives retention
28:10
NVIDIA's dominance is due to its supply chain and CUDA ecosystem.
28:19
Creating a loop between inference and post-training to drive more inference.
33:00
GPU capacity constraints are the main concern
33:53
A clear rubric helps attract and retain the right people
36:44
The on-call culture filters out engineers who avoid pager duty.
38:21
Cheaper inference increases demand, not decreases it
40:41
AI will provide personalized concierge services for everyone
42:34
Thank you for joining us

Chapters

Baseten growth
00:00
Why the app layer wins
01:55
Serving frontier customers
05:57
Open source model mix
07:55
Chinese models and geopolitics
09:21
Custom inference dominates
13:07
Post training acquisition
14:22
When to invest in custom models
17:10
Supply crunch and data centerse
18:35
Longer GPU Contracts
22:25
What Makes a Winner
24:09
Multi Chip Future
26:07
Runtime Roadmap
28:19
Scaling Edge Cases
31:08
Hiring and Leadership
33:48
Operations Pager Culture
36:44
Efficiency Drives Demand
38:19
Concierge Everything Future
40:41
Conclusion
42:34

Transcript

Sarah Guo: Hi, listeners. Today, Elad and I are here with Tuhin Srivastava, the founder and CEO of Baseten, the AI inference cloud. We're here to talk about capacity constraints for AI compute, why inference is the last market, how the workload is changing...