AI news story
Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent.…
Editor's take
Anthropic's Claude Opus 5 has achieved a significant score of 30.2% on the ARC-AGI-3 benchmark, a substantial leap from GPT-5.6 Sol's 7.8%. This development suggests a new level of emergent reasoning capabilities, as Opus 5 reportedly formulated reflection equations autonomously.
This advancement is noteworthy because it signals a potential shift in how we assess and develop AI, moving beyond rote memorization or pattern matching towards more abstract problem-solving. The benchmark's creators expressed surprise at Opus 5's ability to independently derive complex mathematical concepts, a behavior not previously observed in AI models.
Future developments to monitor include independent verification of Opus 5's claimed reasoning process and whether this capability can be replicated or scaled across other LLMs. The long-term impact will depend on whether this emergent property translates to more robust and generalizable intelligence in real-world applications, or if it remains a specific artifact of this particular benchmark.