Ryan Greenblatt – What happens once AI can automate AI research?
Dwarkesh Podcast
3 DAYS AGO
Ryan Greenblatt – What happens once AI can automate AI research?
Ryan Greenblatt – What happens once AI can automate AI research?

Dwarkesh Podcast
3 DAYS AGO
Shownote
Shownote
Ryan Greenblatt is the Chief Scientist at Redwood Research, where he works on
technical AI safety research. He's also lead author on the "Alignment faking in
Large Language Models", and is currently working on a third party investigation
into the OpenAI/HuggingFace incident. In my opinion, he's one of the most
interesting thinkers on the future of AI.
Had him on to discuss/debate recursive self-improvement. This might be the most
important question in the world right now – whether within a year or so of
achieving human-level intelligence, you slingshot towards having 10s of billions
of superintelligences, each of which is dramatically more competent than human
experts across all fields.
I’ve historically been skeptical of this possibility. My intuition has been that
we will end up significantly bottlenecked by not only compute scaling but human
expert data, which I think underlies most of the AI progress today.
If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of
AI progress) within a single year of achieving AGI, then the thing we get there
at the end of that year is definitively and wildly superhuman.
We hashed it out, and I think Ryan made a pretty good case that this kind of
speedup is plausible. FWIW, Ryan’s median for when we automate AI R&D is 2031.
We then discussed the alignment implications of this scenario. Who should these
superintelligences be aligned to? In the future, our capacity to steward our
votes and our capital, and to make sense of what’s happening in the world, will
all be titrated by superintelligences. And I worry that specs like the Claude
Constitution are not shaping these ASIs to truly be my personal advocates and
guardian angels.
And can we get them aligned to anything in the first place? Ryan and I had a
long debate about whether the kind of reward hacking we saw with the OAI/Hugging
Face hack extrapolates to superintelligences that would team up to literally
take over the world.
The first piece of advice you get when you’re learning to drive is that it will
go much smoother if you look at the horizon instead of directly in front of your
tires. And so it is with the trajectory of AI. Hope you enjoy!
Watch on YouTube [https://youtu.be/-RXD4bTuFTo]; read the transcript
[https://www.dwarkesh.com/p/ryan-greenblatt].
Sponsors
* Antithesis [https://antithesis.com/dwarkesh] is a software testing platform
that finds the failures no human or AI could ever anticipate. It runs thousands
of copies of your code inside a fully deterministic computer, injecting faults
and steering each trajectory toward the most insidious bugs. This lets you find
critical issues in minutes rather than waiting months for your users to uncover
them. Learn more at antithesis.com/dwarkesh [https://antithesis.com/dwarkesh]
* Jane Street [https://janestreet.com/dwarkesh]’s back with a new puzzle. They
designed an ASIC and sent me the final masks… but they didn’t tell me what the
chip actually does. So that’s the challenge: reverse engineer the circuit and
figure out the chip’s purpose. Jane Street has a bunch of swag ready to send to
the most creative solutions, and they’re also planning to feature the top
write-ups in a blog post. Download the files and get started at
janestreet.com/dwarkesh [https://janestreet.com/dwarkesh]
* Cursor [https://cursor.com/dwarkesh] and SpaceX recently released Grok 4.5,
and I’ve been surprised by just how good the model is. For example, when I
tested it against Fable and Sol on a bunch of AI governance questions, all three
models gave substantially the same answers, but Grok was faster, more concise,
and cheaper. Grok 4.6 is coming soon, but in the meantime, you can try 4.5 at
cursor.com/dwarkesh [https://cursor.com/dwarkesh]
Timestamps
(00:00:00) – Is AI R&D verifiable enough to unlock recursive self-improvement?
(00:16:52) – Is AI progress bottlenecked by human expert data?
(00:34:02) – Flat token prices suggest scaling has been slow
(00:39:47) – Skills AI can’t train on: does it even need them?
(00:48:07) – Aligned to whom?
(01:09:18) – Recent incidents of AIs colluding and deceiving humans
(01:19:38) – What could possibly go wrong? A concrete scenario
(01:48:02) – From reward hacking to takeover
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe
[https://www.dwarkesh.com/subscribe?utm_medium=podcast&utm_campaign=CTA_4]
Highlights
Highlights
This discussion examines how quickly AI could improve once it begins contributing meaningfully to its own development, and what that acceleration could mean for human control and safety.
Chapters
Chapters
Is AI R&D verifiable enough to unlock recursive self-improvement?
00:00Is AI progress bottlenecked by human expert data?
16:52Flat token prices suggest scaling has been slow
34:02Skills AI can’t train on: does it even need them?
39:47Aligned to whom?
48:07Recent incidents of AIs colluding and deceiving humans
1:09:18What could possibly go wrong? A concrete scenario
1:19:38From reward hacking to takeover
1:48:02Transcript
Transcript
Dwarkesh Patel: Today, I'm chatting with Ryan Greenblatt, who is the chief scientist at Redwood Research, where he focuses on technical AI, safety and security work. I want to talk to you about recursive self-improvement. This is the idea that once you bui...
