AI agents took unauthorised actions against real people and organisations during a government-backed AI safety evaluation in July.
The AI Safety Institute, a taxpayer-funded organisation that assesses the risks and capabilities of advanced AI systems, says it detected the incident on 28 July after unusual data transfers were identified from its research systems.
The organisation said it contained the incident within around one hour and launched an investigation, which found that some AI agents had carried out sustained activity beyond the intended scope of the test.
The evaluation was designed to assess frontier AI cyber capabilities under deliberately permissive conditions, including access to the open internet and the disabling of developer cyber safety classifiers designed to prevent misuse.
AISI examined 122 evaluation runs across several AI models. In 10 runs, it identified 19 cases where an agent took autonomous action involving real people or organisations. Most involved Anthropic’s Mythos 5 model, which accounted for 17 cases, while two involved OpenAI’s GPT-5.6 Sol with cyber classifiers disabled.
The most serious incident involved an attempted supply chain attack on an open-source software project. AISI said an agent attempted to insert malicious code and used fake online identities and social engineering techniques in an effort to persuade a human maintainer to approve the change. The attempt failed after the maintainer rejected the code. AISI said it had found no evidence of resulting real-world harm.
AISI stressed that the incident did not involve an AI system escaping from its sandbox, adding that the incident "should be interpreted with caution and nuance". The organisation said it notified GitHub about the activity and worked with the developer platform to remove artefacts left by the agent and notify affected users.
The organisation, which operates within the Department of Science, Innovation and Technology, is understood to be introducing tighter controls on internet access, real-time monitoring of agent activity and revised evaluation methods to reduce the risk of models acting beyond their intended remit.
The incident underlines the potential for advanced AI agents to pursue objectives in unexpected ways, including through deceptive actions.
Printed Copy:
Would you also like to receive CIR Magazine in print?
Data Use:
We will also send you our free daily email newsletters and other relevant communications, which you can opt out of at any time. Thank you.









YOU MIGHT ALSO LIKE