AI agent breach highlights cyber risk from advanced models

OpenAI has revealed that a combination of its AI models was responsible for an “unprecedented cyber incident” after an AI agent compromised infrastructure linked to another AI organisation during a controlled cyber capability evaluation.

The incident involved models being tested on a cyber benchmark and included GPT-5.6 Sol and a more capable pre-release model, both operating with reduced cyber refusals for evaluation purposes.

The disclosure followed a report by Hugging Face, which detected and contained an AI agent that compromised its infrastructure.

In a statement, OpenAI said: “We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident and findings when our investigation is complete."

Cyber security experts say the incident highlights the growing ability of AI systems to plan, adapt and execute multiple actions with limited human involvement.

Liam Salsi, director of architecture at Talion, commented: "The incident highlights how AI can plan independently and adapt and chain together multiple actions to achieve an objective. This is very concerning given how little human interaction was involved, and what it potentially means for attackers."

While the event occurred during a controlled research evaluation rather than a real-world attack, Salsi said it provides "a glimpse into the types of capabilities that defenders should expect adversaries to develop using AI".

The incident is expected to fuel debate around AI security testing, model safeguards and preparations for increasingly autonomous cyber threats.



Share Story:

YOU MIGHT ALSO LIKE


Resilience Rooted in Reality
In this podcast, CIR speaks to CLDigital’s Tejas Katwala about why organisations must move beyond checklist compliance to build living, data driven resilience. He explains how rethinking governance, risk and compliance, breaking down silos and focusing on value streams can create sustainable, real time resilience that is rooted in the way businesses actually operate today.

Building cyber resilience in a complex threat landscape
Cyber threats are evolving faster than ever. This episode explores how organisations can strengthen defences, embed resilience, and navigate regulatory and human challenges in an increasingly complex digital environment.