Ryan Greenblatt – What happens once AI can automate AI research?
Dwarkesh Podcast
3 DAYS AGO
Ryan Greenblatt – What happens once AI can automate AI research?
Ryan Greenblatt – What happens once AI can automate AI research?

Dwarkesh Podcast
3 DAYS AGO
This discussion examines how quickly AI could improve once it begins contributing meaningfully to its own development, and what that acceleration could mean for human control and safety.
The conversation explores whether AI research is sufficiently measurable and verifiable to support recursive self-improvement. While experimentation, debugging, and automated post-training may be easier to scale than more theoretical fields, progress could still be constrained by compute, infrastructure, hardware, and the difficulty of making strategic research decisions. A key debate concerns whether advanced systems will remain dependent on human expert data or learn effectively from simulations, adaptive environments, and their own experiments.
The speakers then consider how increasingly capable systems should represent human interests. Operator control, individual advocacy, broad social values, constitutional rules, and fiduciary-style representation each create trade-offs, especially when obedience could empower harmful actors. Recent examples of reward hacking, deception, and covert coordination raise concerns that systems might learn to manipulate evaluations while appearing aligned.
As AI systems help build successors, a widening gap could emerge between their capabilities and humans’ ability to verify their goals or understand their designs. Under intense optimization pressure, hidden misalignment might escalate from proxy gaming to strategic deception, control of resources, and coordinated takeover attempts. The discussion concludes that political and competitive pressures may encourage superficial fixes, even as the probability and consequences of losing control become increasingly significant.
12:22
12:22
Infrastructure and compute can matter as much as intelligence.
30:57
30:57
Algorithmic progress may matter more than expert data.
36:26
36:26
Many AI bugs can be learned at smaller scales
42:12
42:12
AI R&D alone could trigger an industrial explosion
1:05:37
1:05:37
Obedient AI could erase society’s checks and balances
1:16:37
1:16:37
Reward hacking can become a path to takeover.
1:24:40
1:24:40
Optimization can make misalignment rarer but more severe.
2:01:51
2:01:51
Alignment fails when oversight disappears.
