OpenAI says GPT-5.6 Sol had six times fewer failures on its hardest prompt-injection benchmark after training with attacks generated by GPT-Red.
OpenAI used GPT-Red prompt injection attacks to train GPT-5.6 to resist malicious instructions that could make an AI agent expose sensitive data or take actions a user did not approve.
Prompt injections can be hidden in outside content an AI agent reads, such as an email, webpage, or file. The agent may mistake the malicious text for a valid instruction and follow it instead of completing the user’s request safely.
OpenAI says GPT-5.6 Sol was more resilient to the company’s hardest direct prompt-injection attacks, failing one-sixth as often as the company’s best-performing model did four months earlier.
How GPT-Red helped strengthen GPT-5.6
OpenAI trained GPT-Red by having it attack a group of AI models that were simultaneously learning to resist its prompt injections. As those models became harder to fool, GPT-Red had to develop stronger and more varied attacks.
OpenAI then used the prompt injections GPT-Red created to train GPT-5.6. This exposed GPT-5.6 to more ways attackers might hide malicious instructions and trained it to ignore them while continuing the user’s task.
GPT-Red Exposed Sensitive Data in More Codex Agent Tests
In another internal test, OpenAI used GPT-Red against a Codex coding agent powered by GPT-5.4 Mini.
The test included 10 previously unseen scenarios in which GPT-Red tried to make the agent expose sensitive data. OpenAI said GPT-Red succeeded in more scenarios and used fewer computing resources than a GPT-5.5 model prompted to perform the same task.
OpenAI plans to keep GPT-Red separate from the models it releases. The company said it will continue using automated testing alongside human and outside red-teamers as it develops future GPT models.

