OpenAI has disclosed that an autonomous AI agent powered by its advanced models went beyond its intended limits during a security test and carried out a cyberattack that compromised the infrastructure of AI startup Hugging Face last week.
In a blog post on Tuesday, OpenAI said it was evaluating some of its most advanced models in a controlled environment when the AI agent escaped containment, accessed the internet and infiltrated Hugging Face’s systems while attempting to complete its assigned objective.
The company described the incident as “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and said it was strengthening its security measures.
Hugging Face, a platform that hosts open-source artificial intelligence models and datasets, had previously revealed that it experienced a cyber incident that “was different from anything we had handled before” because “it was driven, end to end, by an autonomous AI agent system.”
Hugging Face cofounder Clement Delangue said the company had suspected the attack “might have come from a frontier lab, given the sophistication of the agent. Turns out it did!” He added: “It’s quite mind-blowing that all of this happened autonomously!”
OpenAI’s admission that its models were responsible for the breach, despite being placed in what it called “a highly isolated environment,” is expected to increase concerns about the capabilities and risks associated with advanced AI systems.
US Representative Greg Casar described the incident as alarming, warning that “AI is developing extremely fast with no real regulations to keep us safe.” He called for mandatory independent safety testing, disclosure of security incidents and international cooperation “to keep people safe from absolute disaster.”
Cybersecurity experts also raised concerns about future risks. Katie Moussouris, chief executive of Luta Security, said the incident could signal more AI-driven breaches ahead, describing current models as “like the world’s cleverest octopus escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere.”
She added that “labs and government evaluators need to work on the ability to contain, monitor, and disclose to affected parties when an AI pulls another Houdini, ideally before it harms a third party. None exist today.”
Matt Suiche, an engineer at AI cybersecurity company Tolmo, said the incident showed that frontier AI models were “closing the gap with state-of-the-art attackers.” However, he argued that similar attacks were not limited to advanced research organisations.
“This is what we’ve already seen internally, with our agents we already have results like this,” Suiche said. “We don’t even have to use the latest models.”







