AI news story

Fable 5 Beats GPT-5.6 by 15.7 Points — Devs Are Quitting Claude Code for Codex Anyway

Claude Fable 5 leads GPT-5.6 Sol by 15.7 on SWE-Bench Pro — 80.3% against 64.6% in third-party testing. It wins the aggregate…

  • LLMs
  • Source: Towards AI
  • Published: 2026-07-14

Editor's take

Fable 5, a language model developed by Anthropic, has outperformed OpenAI's GPT-5.6 Sol in third-party SWE-Bench Pro evaluations by 15.7 percentage points, achieving an 80.3% success rate in coding tasks compared to GPT-5.6's 64.6%.

This development is significant as SWE-Bench Pro is a widely recognized benchmark for evaluating the coding capabilities of large language models. While Fable 5 demonstrates superior performance on this specific metric, the article notes that developers are continuing to favor OpenAI's Codex for their coding needs, suggesting that benchmark scores alone do not fully capture real-world utility or developer preference. This highlights a potential disconnect between academic evaluation and practical application in the LLM space.

Future developments to monitor include whether Anthropic can leverage Fable 5's benchmark success to influence developer adoption. The persistence of Codex usage despite Fable 5's lead will be telling. It will be crucial to understand if this is due to factors like integration ease, historical developer habit, or specific features not captured by SWE-Bench Pro.