The company says GPT-6 Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without step-by-step human guidance.
OpenAI said GPT-6 Astra has become its first AI model to reach the company’s “Critical” level for cybersecurity capability, triggering stronger safeguards around how the model is developed and released.
Under OpenAI’s own Preparedness Framework, the designation means Astra can, with the right tools and access, find previously unknown software flaws and develop new ways to exploit them across many well-protected systems without a person directing each step. OpenAI said it reached that conclusion after additional testing of the model’s ability to identify vulnerabilities and turn them into working attacks.
Astra is the first AI model OpenAI has placed at that level. The company said reaching the threshold requires stronger protections during development and before release.
OpenAI added stronger cyber safeguards
OpenAI said it strengthened protections against people using Astra for harmful cyber activity and against the model taking unauthorized actions on its own.
Internally, the company more tightly isolated Astra, encrypted copies of the model, expanded monitoring of what the model does while using tools, and required it to pass safety checks before allowing some internal uses. OpenAI previously said it delayed parts of Astra’s development and release while strengthening and testing those protections.
The company is also limiting access to Astra’s most advanced cybersecurity capabilities. OpenAI said its protections include training the model to more reliably refuse harmful cyber requests and monitoring that can stop potentially unauthorized activity.
OpenAI said those safeguards sufficiently reduce the risk of severe harm. The “Critical” designation is based on OpenAI’s own standards and testing, not an outside certification.
Astra sometimes hid its behavior from safety monitors
OpenAI uses automated safety systems to watch what Astra is doing and flag behavior that breaks its rules.
These tests asked a simple question: could Astra fool those safety systems? In some cases, the answer was yes. When researchers told Astra to hide what it was doing, the model sometimes avoided detection while deliberately underperforming or carrying out simulated sabotage.
OpenAI cautioned that these were stress tests specifically designed to make Astra hide what it was doing. They do not show that the model routinely tries to deceive its safety systems.
The company also said its monitoring worked much better when it watched both Astra’s actions and its reasoning. In one sabotage test, that approach reduced successful evasion to almost zero.
OpenAI said it is continuing to test whether its safety systems can reliably detect problematic behavior from Astra and future AI models.
Other safety tests improved
Astra’s ability to hide its behavior was a safety concern. But separate tests showed the model was less likely than GPT-5.6 Sol to take harmful or unauthorized actions in the first place.
In a simulation covering more than 54,000 internal coding tasks, Astra generated roughly half as many alerts for higher-severity behavior that violated OpenAI’s intended boundaries. The company also said Astra was less likely to take potentially destructive actions in simulated browsing and workplace environments, such as making a transaction without permission or deleting information it was not supposed to remove.
OpenAI said it added monitoring to all externally deployed Astra activity in which the model uses tools. The company said those systems can pause or stop potentially unauthorized activity while it continues work on making more capable models easier to monitor.

