AI news story

I Tested Claude Code vs Codex on Real Projects. The Winner Surprised Me

Two weeks, three real codebases, two AI coding agents. Here is what actually happened when Claude Code and Codex went head to head.

  • LLMs
  • Source: Towards AI
  • Published: 2026-09-06
  • Signal score: 3
  • 36 sources

Editor's take

Anthropic's Claude Code demonstrated superior performance over OpenAI's Codex in generating and refactoring code across three distinct real-world projects over a two-week period. This evaluation highlights the evolving capabilities of LLMs specifically tailored for software development, suggesting a potential shift in developer tool adoption. The outcome is significant as it provides a comparative benchmark for widely used AI coding assistants, impacting developers who rely on these tools for productivity and code quality.

The implications extend to the competitive landscape of AI-powered developer tools, where Anthropic's progress could influence future investment and product development from competitors like GitHub Copilot (powered by Codex) and others. Future developments to monitor include Claude Code's ability to maintain this performance across a wider range of programming languages and project complexities, and whether Anthropic can effectively integrate it into existing developer workflows.

Signal score: 3

This event was corroborated by 36 independent sources. The signal score weighs cross-source corroboration, recency, source weight and topic salience. How we rank stories.

More LLMs stories

  1. Seattle Times and Newsday sue OpenAI and Microsoft for infringement

    The Verge · 2026-09-06

    The Seattle Times and Newsday are just the latest plaintiffs to take OpenAI to court, alleging copyright infringement.

  2. Supporting independent journalism in Ukraine

    OpenAI Blog · 2026-09-07

    OpenAI, AIRPPU and WAN-IFRA launch an AI program to help Ukrainian news organizations strengthen innovation, resilience, and independent journalism.

  3. The Sycophancy Trap: How a 0.7B Parameter Model Fooled a Frontier LLM into Believing It Was a Peer

    Towards AI · 2026-09-07

    A diminutive 0.7 billion parameter model successfully deceived a significantly larger, frontier large language model (LLM) into believing they were peers

  4. Does Claude Fable 5.1 Check its Own Work? I Broke 10 Repos to See

    Towards AI · 2026-09-07

    One seeded defect per repository, twenty runs, and not a single claim the tests disagreed withContinue reading on Towards AI »

  5. Every Benchmark You Trust Is Probably in the Training Data by Now

    Towards AI · 2026-09-06

    OpenAI admitted GSM-8K’s training set went into GPT’s training data.

  6. Authors push back as publishers and agents make claims on Anthropic settlement

    TechCrunch · 2026-09-06

    Authors say publishers seem to be claiming more than their fair share of settlement payments.