AI news story
Artificial Analysis overhauls its Intelligence Index after GPT-6 Astra scoring drew skepticism
Artificial Analysis has released version 4.2 of its Intelligence Index, likely in response to criticism that its benchmarks failed to capture GPT-6 Astra's actual progress. Astra now scores four above its predecessor but still trails Anthropic
Editor's take
Artificial Analysis has revised its Intelligence Index methodology following questions about its previous assessment of GPT-6 Astra's capabilities. This adjustment, prompted by skepticism regarding the initial scoring, aims to provide a more accurate reflection of LLM performance.
The recalibration is significant as it impacts how industry observers and developers understand the relative strengths of leading models. The fact that GPT-6 Astra, despite its update, still scores below Anthropic's offerings suggests ongoing competitive dynamics in LLM development, moving beyond simple parameter counts to more nuanced performance evaluations.
Future iterations of the Intelligence Index will be crucial to observe. Specifically, whether this new version can consistently differentiate between incremental improvements and genuinely novel capabilities, and if future model releases from OpenAI, Google, and Meta will fare differently under this updated scoring system, will be telling.
Signal score: 3
This event was corroborated by 31 independent sources. The signal score weighs cross-source corroboration, recency, source weight and topic salience. How we rank stories.
Original reporting
This story summarises reporting published by The Decoder. Read the original article at The Decoder.