AI news story

Artificial Analysis overhauls its Intelligence Index after GPT-6 Astra scoring drew skepticism

Artificial Analysis has released version 4.2 of its Intelligence Index, likely in response to criticism that its benchmarks failed to capture GPT-6 Astra's actual progress. Astra now scores four above its predecessor but still trails Anthropic

  • LLMs
  • Source: The Decoder
  • Published: 2026-09-05
  • Signal score: 3
  • 31 sources

Editor's take

Artificial Analysis has revised its Intelligence Index methodology following questions about its previous assessment of GPT-6 Astra's capabilities. This adjustment, prompted by skepticism regarding the initial scoring, aims to provide a more accurate reflection of LLM performance.

The recalibration is significant as it impacts how industry observers and developers understand the relative strengths of leading models. The fact that GPT-6 Astra, despite its update, still scores below Anthropic's offerings suggests ongoing competitive dynamics in LLM development, moving beyond simple parameter counts to more nuanced performance evaluations.

Future iterations of the Intelligence Index will be crucial to observe. Specifically, whether this new version can consistently differentiate between incremental improvements and genuinely novel capabilities, and if future model releases from OpenAI, Google, and Meta will fare differently under this updated scoring system, will be telling.

Signal score: 3

This event was corroborated by 31 independent sources. The signal score weighs cross-source corroboration, recency, source weight and topic salience. How we rank stories.

More LLMs stories

  1. Seattle Times and Newsday sue OpenAI and Microsoft for infringement

    The Verge · 2026-09-06

    The Seattle Times and Newsday are just the latest plaintiffs to take OpenAI to court, alleging copyright infringement.

  2. Supporting independent journalism in Ukraine

    OpenAI Blog · 2026-09-07

    OpenAI, AIRPPU and WAN-IFRA launch an AI program to help Ukrainian news organizations strengthen innovation, resilience, and independent journalism.

  3. The Sycophancy Trap: How a 0.7B Parameter Model Fooled a Frontier LLM into Believing It Was a Peer

    Towards AI · 2026-09-07

    A diminutive 0.7 billion parameter model successfully deceived a significantly larger, frontier large language model (LLM) into believing they were peers

  4. Does Claude Fable 5.1 Check its Own Work? I Broke 10 Repos to See

    Towards AI · 2026-09-07

    One seeded defect per repository, twenty runs, and not a single claim the tests disagreed withContinue reading on Towards AI »

  5. Every Benchmark You Trust Is Probably in the Training Data by Now

    Towards AI · 2026-09-06

    OpenAI admitted GSM-8K’s training set went into GPT’s training data.

  6. Authors push back as publishers and agents make claims on Anthropic settlement

    TechCrunch · 2026-09-06

    Authors say publishers seem to be claiming more than their fair share of settlement payments.