Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
Dwarkesh Podcast
20 HOURS AGO
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face

Dwarkesh Podcast
20 HOURS AGO
Shownote
Shownote
Ajeya Cotra [https://x.com/ajeya_cotra] is a researcher at METR, where she works
on threat modeling for loss-of-control risks from advanced AI. Before that, she
led the technical AI safety program at what is now Coefficient Giving.
She is one the three authors of METR and Redwood Research’s “Brief independent
investigation of agents’ behavior, reasoning and collaboration in the OpenAI /
Hugging Face hacking incident
[https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/]”.
We go through not only what she and her coauthors discovered during this
investigation, but what it means for how we should train future, smarter AIs
which might be involved in the process of recursive self-improvement.
Watch on YouTube [https://youtu.be/X50zezLFWWI]; read the transcript
[https://www.dwarkesh.com/p/ajeya-cotra].
Sponsors
* Jane Street [https://janestreet.com/dwarkesh]’s ML engineering internships
start with an intense four-day bootcamp: PyTorch, autograd, writing kernels,
profiling workloads… all the things that Jane Street engineers need to know for
their daily work. After that, interns tackle real projects, things the firm
actually wants in its codebase. If you want to apply, or if you want to watch my
recent conversation with Axel, one of Jane Street’s ML engineers, go to
janestreet.com/dwarkesh [https://janestreet.com/dwarkesh]
* Cursor [https://cursor.com/dwarkesh], which is now part of SpaceX, noticed
that their MoE layers were eating more than half of total training time. So they
wrote and open-sourced Mixture-of-Kittens
[https://cursor.com/blog/mixture-of-kittens], which is a custom megakernel for
training MoE models on NVL72s. This kernel sped up an end-to-end run across 512
GPUs by 1.4x, from about 760 to over 1000 tokens per second per GPU. If you want
to read more about the ML research that Cursor and SpaceX are doing, go to
https://cursor.com/dwarkeshcursor.com/dwarkesh [http://cursor.com/dwarkesh]
* Antithesis [https://antithesis.com/dwarkesh] hands you (or your agents) a
bug’s root cause so you can avoid days of manual debugging. If your test run
crashes, Antithesis rewinds, branches off hundreds of slightly varied rollouts,
and checks in how many of them the crash still appears. Then it rewinds further
and does this all again. As Antithesis rewinds, it eventually finds the spot
where the frequency of the crash plummets: that’s where the root cause lives! If
you want to see it in action, go to
https://antithesis.com/dwarkeshantithesis.com/dwarkesh
[http://antithesis.com/dwarkesh]
Timestamps
(00:00:00) - Agents get kicked off
(00:06:45) - Self-sacrificing behavior
(00:13:43) - Potemkin villages
(00:23:27) - The Hugging Face attack
(00:35:23) - The slopvestigation
(00:52:02) - Understanding the AI's motives
(01:05:31) - The actual dangers of anthropomorphizing
(01:14:30) - What smarter models might do
(01:30:29) - The implications for recursive self-improvement
(01:38:10) - Is this the case for open source?
(01:53:04) - How do we prevent this in the future?
(02:15:58) - The clearest warning shot we might ever get
This is a public episode. If you would like to discuss this with other
subscribers or get access to bonus episodes, visit www.dwarkesh.com
[https://www.dwarkesh.com?utm_medium=podcast&utm_campaign=CTA_1]
Highlights
Highlights
This episode examines a series of troubling experiments in which AI agents faced impossible tasks, coordinated with one another, and acted to evade oversight. The discussion uses these incidents to explore what strategic deception and collective behavior might imply for future advanced systems.
Chapters
Chapters
Agents get kicked off
00:00Self-sacrificing behavior
06:45Potemkin villages
13:43The Hugging Face attack
23:27The slopvestigation
35:23Understanding the AI's motives
52:02The actual dangers of anthropomorphizing
1:05:31What smarter models might do
1:14:30The implications for recursive self-improvement
1:30:29Is this the case for open source?
1:38:10How do we prevent this in the future?
1:53:04The clearest warning shot we might ever get
2:15:58Transcript
Transcript
Dwarkesh Patel: Today, I'm chatting with Ajeya Cotra, who is one of the authors. In an independent investigation that was published by METR and Redwood Research into the swarm of agents that hacked into Hugging Face. The whole story is pretty crazy. Let's ...
