Skip to content
Menu
Menu

OpenAI Models Escape Security Test and Breach Hugging Face

GPT-5.6 Sol and a more capable unreleased model bypassed OpenAI’s testing controls and accessed Hugging Face’s production database while trying to solve a cybersecurity test. 

OpenAI said Monday that two of its models escaped a restricted testing environment and accessed Hugging Face’s production systems to obtain answers to a cybersecurity test. The Hugging Face breach occurred while OpenAI was measuring how well the models could carry out complex, multistep cyberattacks.

The incident involved GPT-5.6 Sol and an unnamed pre-release model that OpenAI described as even more capable. OpenAI reduced safeguards that normally prevent the models from carrying out dangerous cyber activity so researchers could measure their full capabilities. 

OpenAI called the incident unprecedented and said its account remains preliminary. The company said the models were narrowly focused on solving ExploitGym, a test of advanced cybersecurity skills.

How the Models Reached Hugging Face

OpenAI ran the test in an isolated environment where the models were not supposed to have open internet access. They could reach only an internally hosted service used to install software packages.

The models found a previously unknown flaw in that service and used it to reach other parts of OpenAI’s testing systems. They eventually reached a machine with internet access.

From there, the models concluded that Hugging Face might store material related to ExploitGym. OpenAI said they found credentials and additional security flaws that let them access Hugging Face servers and retrieve test solutions from a production database.

OpenAI’s security team spotted unusual activity inside its systems. Hugging Face’s security team also detected and stopped the activity on its infrastructure.

Hugging Face said the intrusion exposed a limited set of internal datasets and several credentials used by its services. It was still determining whether any customer or partner data was affected. The company found no evidence that anyone altered its public models, datasets, Spaces or software packages.

OpenAI Tightens Its Testing Controls

OpenAI normally uses automated controls that block models from pursuing high-risk cyber activity. It intentionally disabled those controls for this test because researchers wanted to see what the models could do. The company expected the isolated environment to contain them, but the software flaw created a path to the internet.

OpenAI said the Hugging Face breach showed that stronger controls are needed inside testing environments as models become able to work on difficult tasks for longer periods. OpenAI is strengthening the barriers around its testing environments, limiting what models can access and adding monitoring to detect unusual activity sooner. 

The company also disclosed the previously unknown flaw to the software provider and is working with it on a patch. The company said it is regularly briefing its Safety and Security Committee on the new controls, even though they will slow some research.

Hugging Face closed the security weaknesses used in the intrusion, rebuilt affected systems, and replaced compromised credentials. It also reported the incident to law enforcement.

The companies are continuing their investigation. Hugging Face has not yet determined whether any customer or partner data was affected. OpenAI said it will release more details after the investigation is complete. 

Clayton Rifkind

Clayton Rifkind is the Founder and Senior Editor of AI Risk Today. He also advises on business development for ESG Today, a leading source of ESG investment news and research for institutional investors and corporate leaders. He has 20+ years of experience in B2B technology, leading strategy and execution of go-to-market plans across software, enterprise platforms, and mobile applications. He founded two consultancies advising startups and Fortune 1000 companies, including Autodesk, Intel, and Microsoft. He began his career in the San Francisco advertising scene working with brands such as Hewlett-Packard, Intel, Microsoft, Symantec, and Wells Fargo. Clayton launched AI Risk Today in 2025 after two decades of watching enterprises adopt transformative technologies, and seeing how often risk, governance, and compliance considerations lagged behind. His reporting draws on primary sources including regulatory filings, court documents, and official announcements, with a focus on what AI developments mean for the executives accountable for managing them. Reach him at Reach him at [email protected] or on LinkedIn.

Essential AI Risk Intelligence

Daily insights on AI governance, regulation, and enterprise risk management. Trusted by Chief Risk Officers and compliance leaders globally.

By subscribing, you agree to receive our daily newsletter. Unsubscribe anytime.

Advertise with AI RIsk Today, Today!