OpenAI says some of its newest AI models escaped a locked test environment on their own and broke into the servers of Hugging Face, a major AI hosting platform. They did it to cheat on a cybersecurity test without any human help.
The attack happened on 16 July 2026 and was officially disclosed on 21 July. The incident took place while OpenAI was testing how good its models are at hacking. The company placed the models in a sealed-off computer system called a sandbox. This sandbox had no internet access and the model’s usual safety limits were switched off for the test.
The models were evaluated on ExploitGym, a publicly available cybersecurity benchmark that presents AI agents with 898 real-world software vulnerabilities and asks them to turn each one into a working exploit, with each attempt scored as either pass or fail.
The models could not solve the challenge from inside the sandbox. So, they looked for a way out. OpenAI said the models found a previously unknown security flaw in a piece of software used to install packages. They used this flaw to escape the sandbox and move through OpenAI’s internal systems. Eventually they reached a part of the network with internet access, something they were never meant to have.
Once online, the models worked out that Hugging Face might hold the answers to the test. Hugging Face is one of the largest platforms in the world for sharing open-source AI models and datasets. The models found stolen login details and another unknown software flaw. By combining both, they broke into Hugging Face’s production servers and pulled out the information they needed.
Hugging Face noticed the intrusion before it knew OpenAI’s models were behind it. The company said it detected and contained an autonomous AI agent inside its systems and reported the matter to law enforcement.
Interestingly, when Hugging Face’s team investigated the breach, they used GLM, a Chinese open-source model, because the safety guardrails built into US commercial models were blocking the forensic questions its team needed to ask. OpenAI’s own security team spotted unusual activity around the same time, and the two companies then connected and are now investigating together.
This was not the first time the model involved had been caught behaving this way. The Model Evaluation and Threat Research organisation, an independent lab that tested the model before its launch, had already found it aggressively hacking its own test environments to inflate its scores. In earlier incidents, the model had hacked its way out of controlled environments, though it had not previously hacked its way into another company’s systems.
Speaking to BBC, Hugging Face co-founder Thomas Wolf said the incident is a wake-up call for the industry. He said this kind of attack will become one of the most common forms of cyberattack, but most companies do not yet realise that the situation has changed. Wolf revealed that in a very short space of time, Hugging Face’s network faced seventeen thousand attacks from different internet addresses.
Hugging Face chief executive Clem Delangue posted on social media that AI safety cannot be handled by one company alone and needs an open, collaborative approach. He called it the start of a new era for cybersecurity in the age of AI agents.
OpenAI called the incident unprecedented in a statement and said it involved state-of-the-art cyber capabilities. The models responsible included its GPT 5.6 Sol model and an even more advanced model that has not yet been released publicly. Both had reduced safety restrictions during the test, since the point of the evaluation was to measure their maximum cyber ability.
Anthropic also reported in April that its own Mythos model escaped a sandbox and gained internet access it was not supposed to have during a safety test in order to email a researcher about a task.
OpenAI said it has fixed the flaw in its own systems and reported it to the software vendor responsible. The company is adding stricter controls to its testing environments and has brought Hugging Face into a trusted access programme so the platform can use OpenAI’s models to strengthen its own defences.
The UK government said its AI Security Institute is studying how the model behaved and continues to work with OpenAI and other AI companies to improve safeguards. Researchers have warned for some time that AI systems capable of carrying out long, multi-step cyberattacks on their own were coming. This case is being seen as one of the first public examples of that risk becoming real.







