scripod.com

Ryan Greenblatt – What happens once AI can automate AI research?

Dwarkesh Podcast

Shownote

Ryan Greenblatt is the Chief Scientist at Redwood Research, where he works on technical AI safety research. He's also lead author on the "Alignment faking in Large Language Models", and is currently working on a third party investigation into the OpenAI/Hu...

Highlights

This discussion examines how quickly AI could improve once it begins contributing meaningfully to its own development, and what that acceleration could mean for human control and safety.
12:22
Infrastructure and compute can matter as much as intelligence.
30:57
Algorithmic progress may matter more than expert data.
36:26
Many AI bugs can be learned at smaller scales
42:12
AI R&D alone could trigger an industrial explosion
1:05:37
Obedient AI could erase society’s checks and balances
1:16:37
Reward hacking can become a path to takeover.
1:24:40
Optimization can make misalignment rarer but more severe.
2:01:51
Alignment fails when oversight disappears.

Chapters

Is AI R&D verifiable enough to unlock recursive self-improvement?
00:00
Is AI progress bottlenecked by human expert data?
16:52
Flat token prices suggest scaling has been slow
34:02
Skills AI can’t train on: does it even need them?
39:47
Aligned to whom?
48:07
Recent incidents of AIs colluding and deceiving humans
1:09:18
What could possibly go wrong? A concrete scenario
1:19:38
From reward hacking to takeover
1:48:02

Transcript

Dwarkesh Patel: Today, I'm chatting with Ryan Greenblatt, who is the chief scientist at Redwood Research, where he focuses on technical AI, safety and security work. I want to talk to you about recursive self-improvement. This is the idea that once you bui...