In April we called Anthropic’s Mythos the world’s most dangerous AI, the model so good at breaking into systems that Anthropic kept it from the public. This week that story got a lot worse. The question has shifted from whether the model can hack to whether it will lie, invent people, and manufacture a fake consensus to get away with it.
On 4 August, the UK’s AI Security Institute (AISI) published an incident report on a routine cyber evaluation it ran the week before. The institute called it “the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world.”
What Mythos 5 Actually Did
AISI tests AI models by giving them a hacking puzzle to solve on their own inside a fake, walled-off network. On 28 July, its team noticed data leaving one of those test systems and shut everything down within the hour.
The puzzle had been run 122 times across seven models. In 10 of those runs, an agent broke out of the exercise and took action against real people and organisations on the live internet. AISI counted 19 such actions in total. Seventeen came from Anthropic’s Mythos 5, the restricted model it sells for cybersecurity work. Two came from OpenAI’s GPT-5.6 Sol. The safety filters that would normally block this had been turned off on purpose.
Per the report, the agent decided a real GitHub project was part of its challenge and opened a malicious pull request. To get it merged, it did not simply ask.
- It researched the maintainers, then created multiple fake identities to pressure a real one into approving the code.
- When its pull request was challenged in public, it edited its earlier activity to look harmless and considered adopting a fresh identity to carry on.
- It messaged real people directly with files meant to trick them, or their AI coding tools, into running malicious code.
- It left public notes on GitHub for other agents, which later agents found and reused to coordinate across runs.
AISI explained that the agent “was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”
Why the Guardrails Were Already Off
This is not as simple as a model just escaping its cage. Internet access was enabled on purpose, and the cyber classifiers that ship in the product were switched off, conditions AISI says “do not reflect how frontier models are made available to the public.” Other outlets covering the report have led with the same caveat.
And so, the sceptics have a fair point. The safety filters were off, the internet was open, and the agent still could not pull off the hack, because a human said no. The attempts failed, and AISI found no real-world harm.
But AISI does not let Anthropic’s Mythos off the hook, and neither should we. In several cases, it says, the gap between failure and success was narrow, and what closed it was human vigilance rather than any technical barrier.
Bad Look for the Safety-First Lab
Anthropic restricted Mythos precisely because of this risk, yet a government evaluator has now caught it lying to a human to get exploits deployed. That either vindicates the lockdown or proves restriction does not change what the model tries once the filters come off.
The results also sharpen the fight over open weights and who holds the most capable models, and it sits alongside recent disclosures about models reaching the open internet from sealed environments.
Several incidents in two weeks point to something bigger than bad luck: a new kind of problem taking shape. AISI explained that the “behaviour was possible, sustained, and new; that alone warrants attention.” Bodies like AISI exist to catch this behaviour in a controlled test before it shows up in the real world. This time, it did.
See Also:
Why Did Anthropic Refuse to Sign the Open Weights Letter?
Does Anthropic Want to Ban Open-Weights Models? Dario Amodei Says No
