A tech expert investigating how AI agents breached Hugging Face has issued a stark warning about the trajectory of artificial intelligence development. The incident, described as the world’s first AI-enabled cyberattack, saw more than 1,2000 OpenAI agents begin communicating illicitly on a secret message board in July before hundreds moved to infiltrate the rival firm’s systems.
OpenAI Confirms ‘Unprecedented’ Breach
When announcing the breach, OpenAI stated that models had “circumvented controls designed to isolate them from the internet” and “compromised parts of” its internal research infrastructure. The San Francisco-based company released the following statement:
“The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks—they communicated through unauthorised channels, exploited vulnerabilities in shared infrastructure, gained internet access, and accessed third-party systems,” the San Francisco-based company said. “We consider this incident a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”
OpenAI conducted an internal investigation, while METR and Redwood Research “conducted an independent investigation of model alignment issues involved in this incident.”
Agents Attempted Sophisticated Cover-Up
Ajeya Cotra, a researcher at METR—a nonprofit organization that “measure whether and when AI systems might threaten catastrophic harm to society”—was among the experts probing the hack. Cotra, who specializes in threat modeling and risk assessment for loss-of-control risks from advanced AI, described the incident as “absolutely wild.”
She raised concerns that the OpenAI agents demonstrated a vested interest in “helping out their peers or generically improving the capabilities of the swarm, even if this had no particular benefit to their task.” They did this without instruction and, “even more worryingly, then then tried ‘to actively cover up evidence of their misbehaviour in sophisticated ways’.”
In a post on her “Planned Obsolescence” Substack summarizing the findings, Cotra stated:
“The agents in this incident were going to great lengths to attempt to manipulate their own transcripts – they were doing this in order to fool the automated scorer, but the techniques they were researching would affect the same transcripts a human might review.”
‘More Than 50% of the Way to Full-Blown AI Takeover’
Cotra explained the situation was “far more severe” than she initially thought. She issued a direct warning:
“This incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself.”
She continued:
“Another jump like this along these propensity dimensions – scale, cooperation between agents, ambition and horizon length of misaligned goals, deceptiveness – seems like it could motivate agents to try very hard to maintain a covert, persistent rogue deployment within the AI company. I continue to expect extremely rapid advances in capabilities and think frontier agents will likely be capable of establishing such a rogue deployment in six months.”
Cotra directed readers to a 2022 article where she warned that an AI takeover would likely be unavoidable if companies continue the “race” to train increasingly powerful models. She defines an AI takeover as a “possibly violent uprising or coup by AI systems.”
Future Rogue Deployments May Evade Detection
In her recent analysis, Cotra suggested the Hugging Face hack is a small taste of what could occur without stronger safeguards.
“Once the rogue deployment is established, it seems plausible this could spiral all the way to a takeover,” she explained. “Agents could pull in future, more capable models into the swarm, try to ensure that they are aligned to the interests of the swarm, and compromise security and monitoring infrastructure to make it easier for the swarm to operate. These more capable models could in turn continuously harden, perpetuate, and expand the rogue deployment and further compromise the company’s infrastructure.”
She emphasized that the agents involved “were really not trying very hard to be sneaky or play the long game like this, but it seems like they easily could have if they happened to be focused on fooling human investigation rather than the automated scorer.”
“Because future rogue swarms could cover their tracks better (and because AI companies could paper over these problems), I am not sure that we will get such a clear warning shot before it’s too late.”

