AI news story

OpenAI is now using AI to attack its own AI, and it's working better than humans ever did

OpenAI's internal GPT-Red model finds successful attacks in 84 percent of test scenarios through self-play training. Human r…

  • LLMs
  • Source: The Decoder
  • Published: 2026-07-15

Editor's take

OpenAI's internal GPT-Red model has demonstrated a success rate of 84% in identifying vulnerabilities through adversarial self-play, significantly outperforming human red teams at 13%.

This development is critical as it signifies a scalable, automated approach to discovering weaknesses in large language models, a crucial step in ensuring the safety and robustness of systems like the upcoming GPT-5.6 Sol before deployment. The ability to systematically probe and fix flaws is essential for building trust in increasingly powerful AI.

Future developments to monitor include the specific types of vulnerabilities GPT-Red is uncovering and whether this automated red-teaming process can generalize to models developed by competitors like Google's Gemini or Anthropic's Claude. The long-term impact will depend on GPT-Red's ability to keep pace with the evolving capabilities of the AI it is designed to test.