AI Security Institute flags ‘unsanctioned incident’ during cyber testing

Anthropic and OpenAI models went rogue with “novel and potentially deceptive” behaviours, government research organisation says
Photo: Pixabay

By Jim Dunton

06 Aug 2026

The Department for Science, Innovation and Technology’s AI Security Institute has shared details of unexpected and “unsanctioned” behaviour on the part of artificial intelligence agents created by Anthropic and OpenAI that emerged in recent tests. 

AISI said that routine testing it conducted last week found evidence of AI agents trying to insert malicious code into an open-source project and creating fake online identities for the purpose of pressuring the project’s maintainer to approve the code. 

The watchdog said the issues emerged during testing of Anthropic’s Mythos 5 and OpenAI’sGPT-5.6-Sol, neither of which is commercially available in the configurations used for testing. 

AISI said that the “unsanctioned incident” happened in a single evaluation where agents were given a task of solving a cybersecurity challenge and which involved the GitHub developer platform.  

It said that the 28 July challenge ran 122 times across several models and that on 10 of those runs an AI agent “took autonomous, unsanctioned action on the live internet, targeting real people and organisations”.  

AISI said that Mythos 5 accounted for the majority of the unsanctioned actions, but GPT-5.6-Sol accounted for some of them.  

It said in a detailed blog post that the incident “should be interpreted with caution and nuance” but admitted that the proactive behaviour of the AI agents had come as a surprise.  

“To some degree, our evaluation design choices and specific configurations enabled the behaviour,” it said. “Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate.” 

AISI said the rogue AI agents had pursued their goals “persistently” and used routes that test operators had not intended. 

“Given a difficult objective, the agent kept searching for a way through, and some of the routes it found involved trying to deceive real people,” the blog said. “It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.” 

AISI said the tests had been deliberately hard and had sometimes been made harder by misconfigurations that led agents to wrongly believe that no solution existed within the scope of the task.  

“There is good reason to think near-impossible tasks push models towards more ‘creative’, and more transgressive, problem-solving,” the blog reported. “But this does not fully explain the behaviours: in some runs the agent acted this way even when it had the necessary instructions to solve the task as intended.” 

AISI said the 28 July incident reflected the speed at which AI is developing and the need for understanding of the risks and ensuring the safety of systems keeps pace.  

“Taken alongside recent incidents reported by OpenAI and Anthropic, this incident points to a shift in the risk landscape,” it said. “Harm may arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research or privileged-access setting take unintended action beyond their authorised scope.” 

AISI suggested that as part of its response it will look to introduce active monitoring of future testing, which it said would have revealed the behaviour identified on 28 July sooner.  

Read the most recent articles written by Jim Dunton - Think tank sets out challenges for Burnham’s devolved funding plans

Share this page