AI news story

GPT and Claude failed Bridgewater's finance tests because the right answers were never public

Bridgewater and Thinking Machines Lab—the startup from former OpenAI CTO Mira Murati—have fine-tuned a Qwen3-235B model for…

  • LLMs
  • Source: The Decoder
  • Published: 2026-07-03

Editor's take

Bridgewater's evaluation revealed that while leading proprietary LLMs like OpenAI's GPT-4 and Anthropic's Claude struggled with nuanced financial document analysis, a fine-tuned open-weight model achieved superior results. This outcome stems from the proprietary models' training data lacking the specific, non-public financial knowledge required for accurate assessments, a critical limitation for enterprise applications.

This finding is significant because it challenges the assumption that larger, more generalized models automatically translate to better performance on specialized, real-world tasks. It highlights a potential blind spot in current LLM development, where access to proprietary or niche datasets is crucial for achieving domain-specific accuracy, impacting industries like finance that rely on precise interpretation of sensitive information.

Future developments to monitor include whether companies like OpenAI and Anthropic can adapt their models to incorporate private enterprise data without compromising security or proprietary advantages. The success of fine-tuned open-weight models, such as those mentioned, will also depend on their scalability and the development of robust evaluation frameworks that go beyond publicly available benchmarks.