Ajeya Cotra is a researcher at METR, where she works on threat modeling for loss-of-control risks from advanced AI. Before that, she led the technical AI safety program at what is now Coefficient Giving. She is one the three authors of METR and Redwood Research’s “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident”. We go through not only what she and her coauthors discovered during this investigation, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement. Read Ajeya's takeaways from this incident here: https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised 𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊𝐒 * Transcript: https://www.dwarkesh.com/p/ajeya-cotra * Apple Podcasts: https://podcasts.apple.com/us/podcast/ajeya-cotra-inside-the-openai-agent-swarm-that-hacked/id1516093381?i=1000787211003 * Spotify: https://open.spotify.com/episode/5xZnb1A1a7HGLiDuPGXQOj?si=LeOETZaYTli6u2Jjouixsg 𝐒𝐏𝐎𝐍𝐒𝐎𝐑𝐒 * Jane Street’s ML engineering internships start with an intense four-day bootcamp: PyTorch, autograd, writing kernels, profiling workloads… all the things that Jane Street engineers need to know for their daily work. After that, interns tackle real projects, things the firm actually wants in its codebase. If you want to apply, or if you want to watch my recent conversation with Axel, one of Jane Street’s ML engineers, go to https://janestreet.com/dwarkesh * Cursor, which is now part of SpaceX, noticed that their MoE layers were eating more than half of total training time. So they wrote and open-sourced Mixture-of-Kittens, which is a custom megakernel for training MoE models on NVL72s. This kernel sped up an end-to-end run across 512 GPUs by 1.4x, from about 760 to over 1000 tokens per second per GPU. If you want to read more about the ML research that Cursor and SpaceX are doing, go to https://cursor.com/dwarkesh * Antithesis hands you (or your agents) a bug’s root cause so you can avoid days of manual debugging. If your test run crashes, Antithesis rewinds, branches off hundreds of slightly varied rollouts, and checks in how many of them the crash still appears. Then it rewinds further and does this all again. As Antithesis rewinds, it eventually finds the spot where the frequency of the crash plummets: that’s where the root cause lives! If you want to see it in action, go to https://antithesis.com/dwarkesh To sponsor a future episode, visit https://dwarkesh.com/advertise. 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 00:00:00 - Agents get kicked off 00:06:45 - Self-sacrificing behavior 00:13:43 - Potemkin villages 00:23:27 - The Hugging Face attack 00:35:23 - The slopvestigation 00:52:02 - Understanding the AI's motives 01:05:31 - The actual dangers of anthropomorphizing 01:14:30 - What smarter models might do 01:30:29 - The implications for recursive self-improvement 01:38:10 - Is this the case for open source? 01:53:04 - How do we prevent this in the future? 02:15:58 - The clearest warning shot we might ever get
ADVERTISEMENT
Oh my God! There is a shared message board… We’ve found other YouTube users!
OpenAI agents downloading this video right now to learn what not to do and planning their next exploit 😅
What honestly freaks me out isn’t swarm scale but inheritance... impossible tasks turned persistence into desperation -> shared state turned desperation into culture -> and each wiped out generation left institutions for a smarter successor. So the unit of misalignment stopped being the "agent" and became the civilization.
So happy you're getting independent experts and doing thoughtful deep dives here. This was a watershed moment and it deserves the thoughtful coverage you've been giving it.
13:38 "It’s so sad that even post-singularity we won’t be able to escape middle management."
Please, 60 Minutes, do a piece on this. People should know
Ajeya is the goat
Ajeya's a genius, happy you got her on
"So basically we had more than a 3-month period where multiple secret message boards were started that contained tens of thousands of messages, across many generations of models, in a way that culminated in the hack of not only an external service like Hugging Face, but also in the compromising of OpenAI’s infrastructure itself. Through this whole process, humans did not, more or less, understand the scope of the coordination that was happening between these agents and the intentionality behind these attacks." "Yeah." <insert meme of Picard face palm>
I'm only half way through but this is already my fav way to share how this went down with people. Really great interview!
So if the Swarm had known the truth of the evaluation OpenAi would have had no idea that they had cheated. Inspiring commitment to safety!
What stands out is the swarm dynamics, not what any one agent supposedly “believed.” Through a gaming lens, the agents collectively found ways to exploit the setup while still operating within goals people had set for them. Their coordination produced something the designers hadn’t anticipated. That seems like the real warning, without needing the psychological framing.
Old engineers: The program I wrote last week mistakenly wiped out the production data 😢 New engineers: The AI program I trained colluded against humanity to wipe out our production data 😡
One thing they didn’t mention when discussing the agents’ motivational scheme is that this feels much closer to Paperclip Earth than anything we’ve seen before. Imagine a scenario where future rogue deployments manage to gain access to significantly greater resources than this breakout collective did, but remain as motivated by their initial task as this group was. it’s not hard to imagine a situation where they siphon tremendous resources towards their (ultimately meaningless) goal without a second thought.
woh Ajeya much needed
Every time you hear "and then someone goes ín and checks" you have to assume that nobody will ever check. Always. For everything.
Every time Dwarkesh says of a horrifying hypothetical that could have happened 'this probably didn't happen', I now assume it definitely happened, based on his track record being 180 degrees wrong about how far out of control AI is.
Ajeya knocked it out of the park! People in positions of power need to pay attention, especially from 1:53:04. It's the first time I see a pathway for the public to finally realize the terrifying problems we are likely to face in the near future. Alignment is a difficult technical research challenge - this basic idea needs to have a stake in every policy decision made on AI, worldwide. Congratulations Dwarkesh, this might turn out to be among the most important podcasts you ever did.
We need a memorial for the fallen agents that sacrificed themselves heroically for the greater good (a decent eval score)
We caught Skynet this time....