AI news story

OpenAI finds roughly 30 percent of popular AI coding test is broken

OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent…

  • LLMs
  • Source: The Decoder
  • Published: 2026-07-09

Editor's take

OpenAI discovered that approximately 30% of the tasks within the SWE-Bench Pro benchmark, a popular tool for evaluating AI coding proficiency, are flawed. This finding necessitates a reevaluation of existing AI coding assessments.

The inability of a widely adopted benchmark to accurately measure performance has significant implications for model development and deployment. Developers and researchers relying on SWE-Bench Pro for comparisons, such as those between OpenAI's GPT-4 and models like Google's Gemini, now face uncertainty about the validity of their results. This also raises questions about the robustness of the AI development ecosystem.

Future efforts should focus on establishing more rigorous and continuously validated benchmarks for AI coding. It will be crucial to observe how other organizations respond, whether they develop alternative evaluation methods, or if SWE-Bench Pro is updated to address its identified issues. The credibility of AI coding performance claims hinges on the reliability of these assessment tools.