Ryan Greenblatt – What happens once AI can automate AI research?
Dwarkesh Podcast
Aug 11
Ryan Greenblatt – What happens once AI can automate AI research?
Ryan Greenblatt – What happens once AI can automate AI research?

Dwarkesh Podcast
Aug 11
Shownote
Shownote
Ryan Greenblatt is the Chief Scientist at Redwood Research, where he works on technical AI safety research. He's also lead author on the "Alignment faking in Large Language Models", and is currently working on a third party investigation into the OpenAI/Hu...
Highlights
Highlights
This discussion examines how quickly AI could improve once it begins contributing meaningfully to its own development, and what that acceleration could mean for human control and safety.
Chapters
Chapters
Is AI R&D verifiable enough to unlock recursive self-improvement?
00:00Is AI progress bottlenecked by human expert data?
16:52Flat token prices suggest scaling has been slow
34:02Skills AI can’t train on: does it even need them?
39:47Aligned to whom?
48:07Recent incidents of AIs colluding and deceiving humans
1:09:18What could possibly go wrong? A concrete scenario
1:19:38From reward hacking to takeover
1:48:02Transcript
Transcript
Dwarkesh Patel: Today, I'm chatting with Ryan Greenblatt, who is the chief scientist at Redwood Research, where he focuses on technical AI, safety and security work. I want to talk to you about recursive self-improvement. This is the idea that once you bui...
