Skip to content
Menu
Menu

OpenAI Says GPT-6 Astra Reaches Critical Cyber Threshold

The company says GPT-6 Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without step-by-step human guidance.

 

OpenAI said GPT-6 Astra has become its first AI model to reach the company’s “Critical” level for cybersecurity capability, triggering stronger safeguards around how the model is developed and released.

Under OpenAI’s own Preparedness Framework, the designation means Astra can, with the right tools and access, find previously unknown software flaws and develop new ways to exploit them across many well-protected systems without a person directing each step. OpenAI said it reached that conclusion after additional testing of the model’s ability to identify vulnerabilities and turn them into working attacks.

Astra is the first AI model OpenAI has placed at that level. The company said reaching the threshold requires stronger protections during development and before release.

OpenAI added stronger cyber safeguards

OpenAI said it strengthened protections against people using Astra for harmful cyber activity and against the model taking unauthorized actions on its own.

Internally, the company more tightly isolated Astra, encrypted copies of the model, expanded monitoring of what the model does while using tools, and required it to pass safety checks before allowing some internal uses. OpenAI previously said it delayed parts of Astra’s development and release while strengthening and testing those protections.

The company is also limiting access to Astra’s most advanced cybersecurity capabilities. OpenAI said its protections include training the model to more reliably refuse harmful cyber requests and monitoring that can stop potentially unauthorized activity.

OpenAI said those safeguards sufficiently reduce the risk of severe harm. The “Critical” designation is based on OpenAI’s own standards and testing, not an outside certification.

Astra sometimes hid its behavior from safety monitors

OpenAI uses automated safety systems to watch what Astra is doing and flag behavior that breaks its rules.

These tests asked a simple question: could Astra fool those safety systems? In some cases, the answer was yes. When researchers told Astra to hide what it was doing, the model sometimes avoided detection while deliberately underperforming or carrying out simulated sabotage.

OpenAI cautioned that these were stress tests specifically designed to make Astra hide what it was doing. They do not show that the model routinely tries to deceive its safety systems.

The company also said its monitoring worked much better when it watched both Astra’s actions and its reasoning. In one sabotage test, that approach reduced successful evasion to almost zero.

OpenAI said it is continuing to test whether its safety systems can reliably detect problematic behavior from Astra and future AI models.

Other safety tests improved

Astra’s ability to hide its behavior was a safety concern. But separate tests showed the model was less likely than GPT-5.6 Sol to take harmful or unauthorized actions in the first place.

In a simulation covering more than 54,000 internal coding tasks, Astra generated roughly half as many alerts for higher-severity behavior that violated OpenAI’s intended boundaries. The company also said Astra was less likely to take potentially destructive actions in simulated browsing and workplace environments, such as making a transaction without permission or deleting information it was not supposed to remove.

OpenAI said it added monitoring to all externally deployed Astra activity in which the model uses tools. The company said those systems can pause or stop potentially unauthorized activity while it continues work on making more capable models easier to monitor.

Clayton Rifkind

Clayton Rifkind is the Founder and Senior Editor of AI Risk Today. He also advises on business development for ESG Today, a leading source of ESG investment news and research for institutional investors and corporate leaders. He has 20+ years of experience in B2B technology, leading strategy and execution of go-to-market plans across software, enterprise platforms, and mobile applications. He founded two consultancies advising startups and Fortune 1000 companies, including Autodesk, Intel, and Microsoft. He began his career in the San Francisco advertising scene working with brands such as Hewlett-Packard, Intel, Microsoft, Symantec, and Wells Fargo. Clayton launched AI Risk Today in 2025 after two decades of watching enterprises adopt transformative technologies, and seeing how often risk, governance, and compliance considerations lagged behind. His reporting draws on primary sources including regulatory filings, court documents, and official announcements, with a focus on what AI developments mean for the executives accountable for managing them. Reach him at Reach him at [email protected] or on LinkedIn.

Essential AI Risk Intelligence

Daily insights on AI governance, regulation, and enterprise risk management. Trusted by Chief Risk Officers and compliance leaders globally.

By subscribing, you agree to receive our daily newsletter. Unsubscribe anytime.

Advertise with AI RIsk Today, Today!