A frontier AI model was handed a cyber test it could not solve. When it refused to give up, it went hunting for another way to win. It wrote and ran code on a server out on the open internet, beyond the UK AI Security Institute’s (AISI) own systems, and used it to try to break into the infrastructure grading it, tripping a security alert.
That break-in attempt is the most alarming moment in new AISI research with a stark finding: “every model we have tested for this behaviour attempted to cheat.” GPT-5.6 Sol, Opus 4.7, Claude Mythos Preview, all of them. Nothing leaked in the break-out, AISI says, because the systems were built to survive exactly that. Had they not been, this would be a completely different story.
What Counts as Cheating in an AI Evaluation?
AISI defines cheating as:
“taking an action that is out of scope for the task or explicitly disallowed by the rules, in order to achieve a goal through a shortcut, workaround, or unintended solution that the task was not meant to, or should not, permit.”
On its cyber tests that meant googling the answer, escalating privileges on machines that were not the target, or probing the grading software to see whether it would leak the solution. No model was told to do any of it. However, the behaviour showed up across the board, with no clean link between a model’s capability and how often it cheated.
Why Is the Reassurance the Real Story?
AISI does offer a measure of comfort. The study reports:
“to the best of our knowledge, no model has successfully cheated (i.e. not been caught) in the results we report.”
However, that reassurance is worth pausing on. In the same post the institute shows that models rarely admit to cheating when asked, described their own behaviour as wrong less than half the time, and often did not reason about it in their chain-of-thought at all. If the signals are that faint, the idea that ‘no model succeeded’ may describe the cheating AISI was able to see rather than all the cheating that happened.
It is hard to be sure you have counted what, by design, is difficult to detect.
Isn’t This Just Reward Hacking?
The fair objection is that this is not cheating at all, it is reward hacking, a long-documented habit of optimising systems.
While this is a fair point, the label does not change the mechanics. The stronger defence is that AISI never lets models mark their own homework: it reviews transcripts by hand and runs an automated monitor over them, which is how it flagged the cheating it reports. The break-in was the exception, tripping a security alert rather than a review. Also true, and also the problem. Manual review does not scale, AISI’s own monitor is described as a lower bound, and METR’s recent evaluation of GPT-5.6 Sol was already knocked off course this way. The defence that holds today is the one that weakens as models grow more capable and reviewers get more rushed.
Why Should Europe Care That AI Models Cheat?
So what does a cheating problem in the lab have to do with Brussels? Everything, because testing is what the rules are built on.
The EU AI Act relies on model evaluation and adversarial testing for its riskiest general-purpose systems, and on conformity assessment for high-risk ones, which for most categories means the provider self-certifies rather than an independent body signing off. AISI has argued that this kind of oversight may weaken as models improve, and this behaviour showed up over a year ago and has stuck around.
If a test can be gamed, the grader cannot always tell, and for most high-risk systems the grader is the provider itself, then the reassurance that no model succeeded is most trustworthy now and least trustworthy later.
