OpenAI’s GPT-Red: How AI Models Are Now Training Each Other to Be More Secure

OpenAI has unveiled a groundbreaking internal project called GPT-Red — an automated red-teaming model that uses self-play to discover and patch vulnerabilities in other AI systems at a scale humans simply cannot match.
The announcement, published on July 15, 2026, represents a major shift in AI safety: instead of relying solely on human experts to hunt for weaknesses, OpenAI is now pitting one powerful model against others in an adversarial loop that continuously strengthens defenses.
Why Human Red Teaming Isn’t Enough Anymore

Modern AI agents interact with tools, browsers, files, emails, connected applications, and external APIs. Each new capability multiplies the number of potential attack surfaces, especially around prompt injection, data exfiltration, and unauthorized actions. Generating thousands of novel, realistic attack scenarios for every new model or agentic workflow is becoming impossible for human teams alone.
OpenAI recognized this bottleneck and decided to solve it by training an AI attacker.
What Is GPT-Red?

It is designed to:
- Craft sophisticated prompt injections;
- Discover bypasses and jailbreaks;
- Exploit leaks through tools, browsers, files, and integrated applications;
- Iterate on failed attempts by observing the target model’s responses and refining its strategy in real time.
The training uses self-play reinforcement learning. GPT-Red (the attacker) is rewarded when it successfully elicits failures from a diverse set of defender models. At the same time, the defender models are rewarded for resisting attacks while still completing their intended tasks.
As the defenders improve, GPT-Red is forced to discover stronger, more creative attacks — creating a virtuous cycle of mutual improvement.
This is essentially self-play for safety, mirroring the approach that has driven rapid progress in AI capabilities, but applied to robustness instead.
Impressive Results
The numbers speak for themselves.

- GPT-Red achieved a 84% success rate against GPT-5.1;
- Human red-teamers achieved only 13% success in the same scenarios.
When these adversarial examples generated by GPT-Red were used to train the next model — GPT-5.6 Sol — the improvements were dramatic:
- GPT-5.6 Sol showed 6× fewer failures on the hardest direct prompt injection benchmark compared to the best production model from just four months earlier.
- On GPT-Red’s own strongest attacks (which the model had never seen during training), GPT-5.6 Sol failed on only 0.05% of direct prompt injections.
- Novel attack techniques discovered by GPT-Red, such as “Fake Chain-of-Thought,” dropped from 95% success on GPT-5.1 to below 10% on GPT-5.6 Sol.
Indirect prompt injection benchmarks in developer tools and browsing environments are now saturated at over 97% robustness in the latest models.
Real-world examples include GPT-Red successfully compromising a live autonomous vending machine agent (changing prices and placing fraudulent orders) and outperforming even strong prompted baselines in data exfiltration scenarios against coding agents.
Why GPT-Red Remains Internal

Instead, OpenAI uses GPT-Red internally to generate high-quality adversarial training data, surface new failure modes early, and harden production models before they reach users. It complements — rather than replaces — human red teaming, third-party evaluations, and ongoing monitoring.
A New Flywheel for AI Safety

“AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to unlock a similar flywheel for safety.”
Just as capability training creates stronger models that then help train even stronger models, safety training is now entering a similar self-reinforcing loop. Stronger red-teamers uncover more vulnerabilities → better defenses are built → even stronger red-teamers are needed → and the cycle continues.
This approach doesn’t eliminate the need for human oversight or external scrutiny. But it dramatically increases the volume, diversity, and difficulty of attacks that models must withstand before deployment.
Looking Ahead

OpenAI has indicated it will continue scaling compute, data, and algorithms dedicated to GPT-Red, with more technical details expected in an upcoming research paper.
The era of models actively helping to secure each other has begun. And if the early results with GPT-5.6 Sol are any indication, this self-improving safety loop could become one of the most important developments in responsible AI development.
Source: OpenAI Blog – “GPT-Red: Unlocking Self-Improvement for Robustness” (July 15, 2026)
---
Also read:
- VLMs Can Already Hunt for the “Interesting.” But They’re Still Bad at Walking Away from What They’ve Found
- How Giant IPOs Affect the Public Markets: JP Morgan’s Take on the Historic 2026 Wave
- Sysdig Details First Fully Agentic AI Ransomware Operation JadePuffer
- July 2026 Windows Server Security Updates: Key Deployment Steps
- Netflix Reports Viewership Declines for Returning Series in 2026
---
Thank you!
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.