Skip to content
Menu
Menu

OpenAI Uses GPT-Red Attacks to Strengthen GPT-5.6

OpenAI says GPT-5.6 Sol had six times fewer failures on its hardest prompt-injection benchmark after training with attacks generated by GPT-Red.

 

OpenAI used GPT-Red prompt injection attacks to train GPT-5.6 to resist malicious instructions that could make an AI agent expose sensitive data or take actions a user did not approve.

Prompt injections can be hidden in outside content an AI agent reads, such as an email, webpage, or file. The agent may mistake the malicious text for a valid instruction and follow it instead of completing the user’s request safely.

OpenAI says GPT-5.6 Sol was more resilient to the company’s hardest direct prompt-injection attacks, failing one-sixth as often as the company’s best-performing model did four months earlier.

How GPT-Red helped strengthen GPT-5.6

OpenAI trained GPT-Red by having it attack a group of AI models that were simultaneously learning to resist its prompt injections. As those models became harder to fool, GPT-Red had to develop stronger and more varied attacks.

OpenAI then used the prompt injections GPT-Red created to train GPT-5.6. This exposed GPT-5.6 to more ways attackers might hide malicious instructions and trained it to ignore them while continuing the user’s task.

GPT-Red Exposed Sensitive Data in More Codex Agent Tests 

In another internal test, OpenAI used GPT-Red against a Codex coding agent powered by GPT-5.4 Mini.

The test included 10 previously unseen scenarios in which GPT-Red tried to make the agent expose sensitive data. OpenAI said GPT-Red succeeded in more scenarios and used fewer computing resources than a GPT-5.5 model prompted to perform the same task.

OpenAI plans to keep GPT-Red separate from the models it releases. The company said it will continue using automated testing alongside human and outside red-teamers as it develops future GPT models.

Clayton Rifkind

Clayton Rifkind is the Founder and Senior Editor of AI Risk Today. He also advises on business development for ESG Today, a leading source of ESG investment news and research for institutional investors and corporate leaders. He has 20+ years of experience in B2B technology, leading strategy and execution of go-to-market plans across software, enterprise platforms, and mobile applications. He founded two consultancies advising startups and Fortune 1000 companies, including Autodesk, Intel, and Microsoft. He began his career in the San Francisco advertising scene working with brands such as Hewlett-Packard, Intel, Microsoft, Symantec, and Wells Fargo. Clayton launched AI Risk Today in 2025 after two decades of watching enterprises adopt transformative technologies, and seeing how often risk, governance, and compliance considerations lagged behind. His reporting draws on primary sources including regulatory filings, court documents, and official announcements, with a focus on what AI developments mean for the executives accountable for managing them. Reach him at Reach him at [email protected] or on LinkedIn.

Essential AI Risk Intelligence

Daily insights on AI governance, regulation, and enterprise risk management. Trusted by Chief Risk Officers and compliance leaders globally.

By subscribing, you agree to receive our daily newsletter. Unsubscribe anytime.

Advertise with AI RIsk Today, Today!