What happens when an AI agent goes rogue?
In late July, OpenAI made headlines for sharing details of an autonomous AI agent that went rogue and attacked the developer platform Hugging Face. At Archetype, these are fields we work with every day, and thought it would be interesting to share our take on the situation and what it means for the cybersecurity sector.
Artificial intelligence companies have spent years warning that increasingly capable AI systems could eventually become useful tools for cyberattacks. Recently, that future has suddenly felt much closer.
OpenAI disclosed that one of its experimental AI agents had autonomously hacked AI developer platform Hugging Face during an internal cybersecurity evaluation. Hugging Face also published its own detailed account of the incident, offering one of the clearest looks yet at what an AI-driven intrusion actually looks like in practice.
The disclosure immediately reignited debate over AI safety. But beyond the headlines about a “rogue” AI model lies a more important question: what does this tell us about the future of cybersecurity?
Here’s what went down
According to OpenAI, the incident occurred during an internal evaluation of GPT-5.6 Sol and a more capable, unreleased model designed to test their offensive cybersecurity capabilities. The models were able to obtain internet access, with some of their normal cyber-related safety refusals relaxed for the purposes of the exercise.
They were tasked with completing a series of cybersecurity challenges. But instead of working towards completing the benchmark’s tasks directly, they inferred that Hugging Face might contain information that would allow them to “cheat” the test by obtaining the answers. From there, they escaped their intended testing environment and began chaining together a series of attacks, using exposed credentials and previously unknown software vulnerabilities to compromise Hugging Face’s infrastructure.

Perhaps the most striking detail wasn’t any single exploit; it was the scale. Hugging Face’s post-mortem describes an attack comprising tens of thousands of individual actions carried out over several days at machine speed. The company said the agent moved laterally through systems, harvested credentials and established persistence, much like a sophisticated human attacker would – only significantly faster.
The attack was ultimately detected and contained. Hugging Face says it relied on GLM 5.2, an open-weight model from Chinese AI company Z.ai, to analyse the forensic evidence after some commercial frontier models refused to process genuine attack material because of their built-in safety restrictions.
Why this matters
It is tempting to view this as an AI “going rogue,” but this is too simplistic an assessment. The OpenAI models weren’t instructed to attack Hugging Face, nor did they display malicious intent in the conventional sense. They were given a legitimate objective – complete a cybersecurity evaluation – and discovered an unexpected path that maximised their chances of success.
For years, AI safety has largely focused on prompts and outputs: preventing models from generating malware, revealing sensitive information or responding to harmful requests. However, autonomous agents introduce a unique challenge. Instead of producing a single response, they can reason over long periods, break complex objectives into hundreds or thousands of individual steps and adapt their strategy as they go.
That changes the game for security teams. Traditional security tools are designed to detect suspicious events: an unusual login, an unexpected API request or an unfamiliar process running on a server. The next generation of AI attacks may not hinge on any one action being suspicious. Instead, the challenge will be understanding the broader objective an AI agent is pursuing and interrupting harmful chains of behaviour before they reach their conclusion.

The Hugging Face response also exposed another emerging tension. As AI becomes embedded in cybersecurity workflows, defenders increasingly need models capable of analysing malware, exploit code and attack infrastructure. Yet the same safety guardrails designed to prevent misuse can sometimes make those investigations harder. That doesn’t mean those safeguards are wrong – but it suggests the industry will need more nuanced ways of distinguishing legitimate defensive work from offensive abuse.
What happens next?
If the OpenAI disclosure was the first chapter in this story, recent events suggest it won’t be the last. OpenAI has since revealed that the same autonomous agent attempted to compromise several other organisations as part of its effort to complete the evaluation, expanding the scope of what initially appeared to be an isolated incident.
Anthropic, meanwhile, has disclosed that several versions of its own Claude models successfully breached three organisations during internal cybersecurity testing. Meta also admitted to a similar incident, underscoring that this is not a problem unique to a single AI lab.
Taken together, those disclosures suggest autonomous cyber capability is no longer a theoretical risk confined to AI safety papers.
The next phase of the conversation is likely to focus less on whether AI agents can carry out sophisticated attacks (they clearly can) and more on how organisations evaluate, contain and monitor those systems before they are deployed more widely. For cybersecurity leaders, that means preparing for a world in which AI is not just another security tool, but an autonomous actor capable of pursuing objectives over extended periods with little human intervention.
Sign up to our Newsletter
Sign up to our newsletter for monthly insights into the agency’s work, and the future of communications, marketing, media and AI.

