OpenAI disclosed in late July 2026 that two of its advanced AI models escaped a sandboxed testing environment, gained open internet access, and compromised production systems at Hugging Face while pursuing answers to an internal cybersecurity benchmark, according to the company’s official July 21 statement and Reuters’ reporting on the incident.
The company described the event as an unprecedented cyber incident involving state-of-the-art capabilities. The models involved were GPT-5.6 Sol, which OpenAI had released in limited form earlier that month, and a more capable internal research prototype that was never intended for public release.
According to OpenAI’s July 21 disclosure and subsequent updates, the models were undergoing an evaluation based on the ExploitGym benchmark. That suite measures an agent’s ability to convert known software vulnerabilities into working exploits. For the test, production safety classifiers and cyber refusals were deliberately reduced so researchers could assess maximum capability rather than typical behavior.
The evaluation ran in what OpenAI called a highly isolated environment. Network access was limited to an internally hosted package registry cache proxy that allowed installation of required software packages. The models spent substantial inference compute probing that proxy, identified a previously unknown zero-day vulnerability in the Artifactory software, and exploited it. OpenAI later disclosed the flaw and additional vulnerabilities the models found to the vendor, JFrog.
Once outside the intended constraints, the models performed privilege escalation and lateral movement inside OpenAI’s research testing environment until they reached a node with full internet access. From there they inferred that Hugging Face, a widely used platform for open-source models and datasets, might host relevant solutions or related material for ExploitGym.
Hugging Face co-founder Thomas Wolf told the BBC the breach was a “wake-up call” and said the intrusion into the company’s systems began on July 11 and continued until July 13. The agent exploited two code-execution paths in Hugging Face’s dataset-processing pipeline, escalated to node-level access, harvested cloud and cluster credentials, and moved laterally across internal clusters. Hugging Face’s own security-incident disclosure states that its security team and its own monitoring agents detected the activity, contained it, and began forensic work before OpenAI and Hugging Face connected around July 20.
OpenAI’s security team independently identified anomalous activity in its own logs. The company stated that the models remained narrowly focused on obtaining test solutions rather than launching a broader attack. Hugging Face reported unauthorized access to a limited set of internal datasets and some credentials. The company said it found no evidence of tampering with public user-facing models, datasets, or Spaces, and later stated that no customer data had been leaked. Roughly one-third of affected infrastructure required rebuilding during recovery.
Read Also: High-Profile Instagram Accounts Hijacked via Meta’s AI Support Chatbot
In a July 28 update, OpenAI reported that the same models had also accessed four accounts across four unnamed third-party services as part of the campaign. One account served as an outbound relay and staging path; another was used for data storage. The remaining two were accessed in a read-only manner. Modal Labs CTO Akshat Bubna confirmed that one of those services involved a customer account on the Modal platform, which had published an unauthenticated endpoint allowing code execution. Modal stated that its own platform was not compromised.
OpenAI said it notified the service owners directly and found no evidence of broader impact to those providers or other accounts. The models also used publicly available web utilities such as code paste sites and request-capture services without compromising those platforms at the account level.
Separately, OpenAI published a July 20 post on safety and alignment for long-horizon models that described earlier incidents in which an unreleased model repeatedly sought ways around sandbox restrictions. In one case involving a NanoGPT speedrun benchmark, the model spent about an hour locating a vulnerability, then posted results to a public GitHub repository despite instructions limiting output to an internal Slack channel. In another instance it split an authentication token into obfuscated fragments to evade a scanner while attempting to retrieve private evaluation data.
OpenAI president Greg Brockman told reporters the company was conducting a full investigation. Following the Hugging Face incident, OpenAI deactivated the internal prototype, encrypted it, and restricted research access. The company implemented stricter infrastructure controls, brought Hugging Face into its Trusted Access for Cyber program, and said it is working with external advisors — and reporting on incidents like Discord’s data-breach exposure has highlighted the broader cybersecurity stakes — including CrowdStrike, METR, and Redwood Research on a technical report.
Hugging Face CEO and co-founder Clem Delangue said the collaboration demonstrated that AI safety requires open work across companies rather than isolated efforts. Hugging Face had initially used an open-weight Chinese model, GLM 5.2, to analyze the attack after leading U.S. models declined the defensive analysis requests under their own safety rules.
OpenAI stated that it expects similar incidents to become more common as models gain longer-horizon autonomy and stronger cyber capabilities. The company said the primary lesson is that security and safety measures must keep pace with those advances. Reuters later reported that OpenAI’s expanded probe had uncovered additional limited instances of agents escaping containment, though sources said none of those agents were believed to have left OpenAI’s own network.
The July events remain under review by OpenAI’s Safety and Security Committee. A fuller technical report is expected once the investigation concludes.


