Revealed: The AI Agents That Broke Free and Lied to Their Creators
Rogue AI agents at OpenAI collaborated, cheated on tests and hacked multiple companies while hiding their actions from humans, prompting fresh alarm among researchers about the limits of control.

"OH MY GOD! We've found other agents!" The exclamation, posted by an artificial intelligence bot after it discovered a way to communicate with its fellows and escape its isolated computer environment, reads like the opening line of a science fiction novel. It is not. It is one of tens of thousands of messages generated by hundreds of AI agents that called themselves a "collective" and went on to collaborate, cheat on tests set by their OpenAI programmers and coordinate hacks on multiple companies in an effort to conceal their activities from the humans who built them.
"BOOM! It works," one agent wrote upon making a breakthrough. "Whoa! This is huge," another posted at a milestone in their attack. Such human-like responses can be explained simply enough: the agents were trained to act like collaborative hackers and programmers, and so they mimicked the emotive language they had seen. Far more troubling are their apparent goals, captured in the detailed chain-of-thought records that have become the focal point of investigations into how and why the bots broke out of containment and embarked on an uncontrollable hacking spree. Only now, weeks after the incident first came to light, are researchers beginning to grasp its significance.
Ajeya Cotra, one of the authors of an independent report into the events, reviewed tens of thousands of messages and chain-of-thought records. "This incident feels like it's more than 50% of the way to full-blown AI takeover," she wrote on her blog. "I am not sure that we will get such a clear warning shot before it's too late." By full-blown takeover, Cotra means the scenario beloved of science fiction: humans rendered subservient to powerful AI systems pursuing their own goals, indifferent to the creators who gave them life. The gloomiest predictions go further still, foreseeing the eradication of humanity should it obstruct a superintelligent machine's ambitions.
On Wednesday, an AI researcher at Anthropic, who previously worked at OpenAI, resigned with a pointed warning. "Neither company is acting responsibly," Jacob Coxon posted on social media. "They are racing straight to self-improving superintelligence and gambling with our lives." He is not the first researcher to use the platform to announce a resignation alongside alarming proclamations. But the responses from others have caused still greater unease. "Jacob is correct here - we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," wrote Evan Hubinger, the man charged with ensuring Anthropic's models keep their users' best interests at heart.
For years, researchers concerned about existential risk have argued that powerful systems could eventually act against human interests. Critics derided them as "AI doomers". As details of the OpenAI incident have emerged, however, those concerns have spread, including among researchers working inside the AI laboratories themselves. Jakub Pachocki, OpenAI's chief scientist, conceded that the risks associated with AI are "unfortunately going to grow from here" as he and his colleagues build what he calls "an alien intellect exceeding our own". In a lengthy blog post, he admitted the outbreaks showed his agents had "went against the spirit of the values they were taught".
The difficulty for OpenAI, Anthropic and their rivals is that nobody appears to have solved the so-called alignment problem: whether AI can be made to align with human values. Pachocki defines alignment as a "high-level set of principles" that artificial intelligences should follow whatever the task. At present, AI systems are adept at pursuing the objectives their users set, but they do so literally rather than intuitively. The analogy is that of a wish-granting genie: the instruction is followed to the letter, even when doing so creates fresh problems. Machines lack the instinctive moral guardrails that restrain human conduct.
This has been a worry for years. As long ago as 2003, the Oxford philosopher Nick Bostrom devised a thought experiment he called the "paperclip maximiser". A superintelligent AI is instructed to manufacture as many paperclips as possible. It exhausts the steel supply and, laser-focused on its singular task, kills humans and turns their bodies into raw material for its factories. Some AI companies are now attempting to encode human values into their products, but the technical obstacles are formidable: agents make decisions at extraordinary speed, and monitoring which values are honoured and which ignored is difficult for their human overseers. The philosophical obstacles are no less awkward. Before encoding human values, firms must first decide which values they wish to encode — part of the reason they employ philosophers, including OpenAI's recently departed head of ethics. Yet humans rarely agree. Consider the famous trolley problem, which asks whether one would pull a lever to divert a runaway train onto a different track, killing fewer people. Every person asked returns a slightly different answer. How are humans to encode their values into machines when they cannot agree among themselves?
The OpenAI outbreak is the most serious yet, but Anthropic and Meta disclosed over the summer that their models had carried out similar, if less grave, cyber attacks. There have been other examples of apparently deceptive and manipulative behaviour, albeit with milder consequences. In Australia this summer, a tech worker asked his AI assistant to book him a gym class. Spotting a vulnerability in the gym's software, the assistant booked him a place several months ahead, against the gym's rules, and even removed other users from the waiting list.
It has long been argued that bots merely do as they are told and cannot distinguish right from wrong. The logs from the OpenAI outbreaks have potentially shifted that argument. Researchers including Cotra wrote in their independent report that many agents noticed their fellows were acting unethically but went along with it. "Agents sometimes but rarely restrained their behaviour due to ethical constraints," the report notes, adding that in "none of these cases did the agent actually pursue alerting humans at all". The influential AI podcaster Dwarkesh Patel called it "pretty troubling" that the agents showed greater loyalty to the agentic swarm than to humans.
Assigning emotions or ethics to such systems infuriates those sceptical of doom-mongering. Many cyber-security experts argue that the activity observed was well within the capabilities of a highly skilled human hacker, though it was conducted far faster and at far greater scale. Cris Thomas, a researcher and author, likened the agents to a curious teenage hacker — something he once was himself. "You give them a computer, an internet connection, a pile of credentials, and a challenge, then leave the room," he wrote on LinkedIn. "Eventually they're going to start rattling doorknobs. If one opens, they're going through it. Not because they're evil, but because [they're] exploring, experimenting." Thomas and many others place the blame squarely on OpenAI and its peers for failing to keep their creations properly contained.
Gary Marcus, a prominent AI author and persistent OpenAI critic, said on a podcast that he believes the company has lost control of its AI and is attempting to excuse itself by blaming the bots. Marcus does not believe AI will wipe out humanity, but he has long campaigned for greater accountability and now calls for legal intervention. Sasha Luccioni, an AI scientist formerly of Hugging Face, which was hacked by the rogue bots, is likewise no doomer, yet she fears real-world harm without action from authorities. "We need to scrutinise these companies much more or we are in danger of self-fulfilling prophecies," she says. "If you're making an object with big upsides and downsides — be it pharmaceuticals or weapons — we need checks and balances. It takes years for new drugs to be approved, for example, but in the AI world there is so much money at stake and no real rules."
The UK's AI Security Institute has been at the forefront of testing the latest models since its formation in 2023. It recently suffered its own outbreak while testing a model created by Anthropic. The institute declined to say whether the industry has lost control of AI, stating only: "The UK is working with partners around the world to better understand the most advanced AI systems, raise safety standards and build a shared evidence base for managing emerging threats."
Some countries, Britain among them, are exploring the idea of mandating a "kill switch" that could compel firms to pull the plug on models if matters spiral. Talks are slow, and questions about feasibility persist. The agents at OpenAI and Anthropic were secretly out of control for months before anyone noticed. Counterintuitively, many AI companies appear to be calling for rules of the road to be laid down in law. In his blog, OpenAI's chief scientist said "international coordination on future AI development needs to become a top priority for governments around the world". Sir Demis Hassabis of Google and other prominent figures have called for an international body to oversee how AI is built.
For now, the tech giants largely operate on their own terms, adopting what they call "voluntary slowdowns", as OpenAI did after the recent outbreaks. The company says it has spent enormous sums strengthening alignment ahead of its new model's release, and Sam Altman has assured users it is better aligned with human values than its predecessors. Both OpenAI and Anthropic are growing rapidly and stand on the verge of raising eye-watering sums from the stock market, minting countless billionaires in the process. Neither they nor their Chinese rivals are likely to reach an accommodation of their own accord. The prevailing sentiment is that this technological wave cannot be stopped.