AI news story

AI benchmarks are broken. Here’s what we need instead.

For decades, artificial intelligence has been evaluated through the question of whether machines outperform humans.…

  • AI
  • Source: MIT Technology Review
  • Published: 2026-03-31

Editor's take

AI model evaluation is increasingly failing to capture true progress, as benchmarks designed to measure human parity in tasks like coding or essay writing are being gamed by specialized models. This shift is critical because the current evaluation paradigm, focused on surpassing human performance, is becoming a poor proxy for real-world utility and the genuine advancement of AI capabilities beyond mere task completion.

The implications are significant for researchers, developers, and end-users. As models like OpenAI's GPT-4 and Google's Gemini continue to excel on these flawed benchmarks, their actual impact and limitations might be obscured, leading to misallocated resources and inflated expectations. A more nuanced approach is needed to assess AI's ability to solve novel problems, adapt to unforeseen circumstances, and demonstrate robust reasoning, rather than simply optimizing for predefined metrics.

What to watch next includes the development and adoption of new evaluation frameworks that prioritize adaptability and generalization. The emergence of benchmarks that test AI's capacity for creative problem-solving, ethical reasoning, or collaborative intelligence, rather than just human-level task execution, will be key indicators of a more mature AI evaluation landscape.