Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]
Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]
Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]
This episode examines how AI infrastructure may evolve as systems shift from instant responses to sustained, autonomous work. Neil Movva explains the technical and economic challenges involved in making that future practical.
Movva describes Sail’s goal of building a highly efficient inference platform for long-running AI agents, where throughput, utilization, and energy costs matter more than minimal latency. He argues that persistent agents could support continuous research, cybersecurity, and other complex tasks, but only if token generation becomes dramatically cheaper.
The conversation explores the full intelligence stack: model architecture, software kernels, memory, accelerators, networking, data centers, and power. GPUs remain powerful and flexible, but specialized chips such as Cerebras- and Groq-like systems may offer advantages for particular workloads. Key bottlenecks include memory bandwidth, KV-cache storage, chip supply, advanced packaging, electricity, and the difficulty of scaling centralized facilities.
Sail’s strategy is to combine heterogeneous hardware, underused compute, distributed sites, and intermittent energy sources, while improving orchestration and compression. Movva expects open and closed models to coexist, with falling inference costs enabling more customizable agents. He also questions whether frontier labs can preserve their current premium as open models, generated data, and distillation narrow their lead, and suggests that future AI hardware will ultimately be shaped by the models and workloads it must serve.
00:00
00:00
AI inference can become dramatically cheaper
02:22
02:22
Cheaper inference could unlock longer-running AI agents
03:22
03:22
Sail aims to make AI inference dramatically cheaper
07:44
07:44
Background workloads could dominate token consumption
11:23
11:23
Cheap inference makes long-horizon AI practical
17:30
17:30
Interactive AI demands speed, not just throughput
22:28
22:28
Interconnects can outweigh raw compute efficiency
29:49
29:49
Long-context AI turns KV cache into a memory bottleneck.
35:58
35:58
Overfitting reveals patterns that enable generalization
39:12
39:12
Software efficiency may not remain a durable moat
44:06
44:06
No chip is bad at the right price
50:31
50:31
Smaller sites make large-scale AI more practical.
53:40
53:40
AI agents can trade uptime for lower costs
59:07
59:07
AI inference still leaves enormous efficiency gains untapped
1:06:03
1:06:03
Technological leads last only while they remain difficult to copy
1:09:22
1:09:22
AI must become millions of times cheaper to scale
1:10:37
1:10:37
Model architecture will shape AI infrastructure
1:12:50
1:12:50
NVIDIA’s biggest vulnerability may be limited HBM supply
1:14:39
1:14:39
Understanding the full stack creates lasting advantage

![Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]](https://image.scripod.com/https://megaphone.imgix.net/podcasts/d56fad4c-876e-11f1-9151-5f1834daf707/image/de9954428adab72e672734106004bdcb.jpg?ixlib=rails-4.3.1&max-w=3000&max-h=3000&fit=crop&auto=format,compress)