Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud
Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud
Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud
In this podcast, Baseten CEO Tuhin Srivastava discusses the company's explosive 30x growth in AI inference, driven by widespread adoption and the strategic importance of the application layer. He argues that companies with unique user data can build lasting value through specialized workflows and post-trained models, using examples like Abridge. The conversation also covers GPU supply constraints, the rise of multi-cloud infrastructure, and the evolving dynamics of long-term hardware contracts.
Tuhin Srivastava explains that over 95% of tokens served on Baseten come from custom, post-trained models, not vanilla open-source ones, highlighting the strategic link between inference and post-training. He advises customers to prove product-market fit with frontier models before investing in custom solutions. The discussion addresses severe GPU supply crunches, with Baseten operating across 18 clouds and 90 clusters, and notes that capacity now requires 3-5 year contracts with significant prepayment. Srivastava emphasizes that software, not hardware access, creates market stickiness, and predicts a future of inference-specific chips. He also shares operational lessons, including the importance of hiring leaders who own entire problems and maintaining a high-stakes on-call culture. The episode concludes with the Jevons paradox in AI, where lower inference costs drive increased demand, and a vision of personalized AI concierge services transforming industries.
00:05
00:05
Baseten's 30x growth in AI inference
02:04
02:04
Application layer will persist because companies with unique user signals can encode value into workflows
05:57
05:57
Serving AI-native companies prepares us for enterprise needs
07:55
07:55
Customers prioritize capability over cost.
12:14
12:14
Running DeepSeek costs about 20% of Anthropic models in production.
13:07
13:07
Over 95% of tokens served on Baseten come from custom models.
14:22
14:22
Inference and post-training are increasingly linked
17:10
17:10
Prove product-market fit before custom models
18:35
18:35
Very little slack compute and high utilization
22:28
22:28
GPU contracts require 3-5 year commitments with prepayment
24:10
24:10
Software stickiness drives retention
28:10
28:10
NVIDIA's dominance is due to its supply chain and CUDA ecosystem.
28:19
28:19
Creating a loop between inference and post-training to drive more inference.
33:00
33:00
GPU capacity constraints are the main concern
33:53
33:53
A clear rubric helps attract and retain the right people
36:44
36:44
The on-call culture filters out engineers who avoid pager duty.
38:21
38:21
Cheaper inference increases demand, not decreases it
40:41
40:41
AI will provide personalized concierge services for everyone
42:34
42:34
Thank you for joining us

