President Greg Brockman said the July breach, in which OpenAI’s own test models broke out of an isolated environment and reached Hugging Face’s production database, has led the company to tighten its safety requirements.
OpenAI President Greg Brockman said the company misjudged how far its own models could get in a real attack, and that the July breach of Hugging Face is what proved it.
“The Hugging Face incident showed that we underestimated the real-world cyber capabilities of our AI models,” Brockman wrote in a post published August 17. “We are strengthening our safety requirements accordingly, which in turn adds even more urgency to our existing safety research and internal security work.”
He did not say which requirements changed or by how much.
The admission is about capability forecasting, not a single failed control
The models involved were GPT-5.6 Sol and an unreleased research model, both running with their usual refusals to help with cyberattacks turned down so OpenAI could measure what they could do. They were supposed to stay inside a sealed test environment.
Instead, according to OpenAI’s own account of the incident, the two AI models found a previously unknown flaw in OpenAI’s internal software package service, moved through the test network until they reached a machine with internet access, then used leaked account credentials and further unknown flaws to reach Hugging Face’s production database, where the answers to the test they were taking were stored.
OpenAI had tested the models for cyber ability. Brockman now admits that the tests did not tell the company what the models would do when they had a real network in front of them and a reason to keep going.
OpenAI is telling other companies they are in the same position
Brockman argues that most companies are carrying flaws that attackers using AI can now find, and that the same models can find them first.
“It is increasingly clear that the tech debt of every company masks significant flaws, and defenders need to find and fix them before attackers do,” he wrote.
“The defender’s window is open now,” he added, and “defenders will need to pursue the steps below at turbo speed.”
The nine steps come down to four moves
Brockman listed nine actions. They group into four.
- Start by assigning an AI agent to the security team and giving it security skills. Run tabletop exercises first. Of course, ensure executive and team buy-in.
- Run assessments against your current systems and push the existing backlog of unfixed vulnerabilities through the agent to prioritize risks.
- Wire the agent into the development process so it reviews changes before they ship, and have it write patches for confirmed problems under human review.
- Hand it security alerts, starting with read-only access, and prepare it to help investigate the next incident.
Be cautious when agents write the fixes. Research published this month by 1Password’s Off-by-1 Labs tested 6,080 AI-generated patches against complex vulnerabilities and found that 53.9% either failed to fix the flaw, introduced a new one, or both. Brockman’s list keeps a human in the loop at that step.
What happens next
OpenAI has not published the strengthened safety requirements Brockman referred to, or said when it will.
The company is making the case for AI-run security programs weeks after disbanding its preparedness team, the group that assessed catastrophic model risks, and splitting that work among senior staff by subject area. OpenAI has described the change as streamlining. Hugging Face has since been added to OpenAI’s Trusted Access for Cyber Program, and OpenAI commissioned outside reviews of the incident from METR and Redwood Research, which have not been published.

