AI news story
Anthropic Secretly Nerfed Claude for Six Weeks. Then They Admitted It.
Reasoning effort dropped from HIGH to MEDIUM on March 4. A caching bug deleted reasoning history on March 26. Benchmark accur…
Editor's take
Anthropic appears to have inadvertently reduced Claude's reasoning capabilities for an extended period due to a caching bug, a fact only revealed after the issue was identified and partially addressed. This incident highlights the fragility of complex LLM architectures and the critical need for robust internal monitoring and transparent communication. The six-week duration of the "nerfed" state, during which benchmark accuracy reportedly declined by 18%, suggests that even sophisticated models can suffer significant performance degradation from seemingly minor technical glitches, impacting users who rely on consistent AI behavior.
The implications extend beyond Anthropic, raising questions about the reliability and auditability of proprietary LLMs across the industry. Users and developers are left to ponder how frequently such performance anomalies might occur undetected in other systems. Future developments to watch include Anthropic's detailed post-mortem analysis and any subsequent architectural or testing changes they implement to prevent recurrence. The industry will also be observing whether competitors face similar unannounced performance shifts and how they choose to disclose them, if at all.