AI news story
Claude vs GPT-5.6 vs Gemini vs DeepSeek Prompt Caching: I Turned It Off and Saved 20%
I turned prompt caching off on a Claude Opus 5 assistant last week and the input bill dropped 20% — from $4.61 to $3.69 for the same 41…
Editor's take
Disabling prompt caching for Anthropic's Claude Opus 5 demonstrably reduced API input costs by 20% in a specific usage scenario. This finding directly impacts developers and businesses relying on large language model APIs, particularly those with high-volume, repetitive query patterns. The current economic pressures on AI deployment make such cost optimizations significant, suggesting prompt caching, while potentially beneficial for latency, might be an underappreciated cost driver.
The observed cost reduction raises questions about the default configurations and transparency of prompt caching mechanisms across major LLM providers like OpenAI (GPT series) and Google (Gemini). Future developments to monitor include whether other developers observe similar savings, if providers offer more granular control over caching, or if alternative caching strategies emerge that balance performance and cost more effectively. The long-term viability of certain AI applications hinges on such economic efficiencies.
Signal score: 5
This event was corroborated by 9 independent sources. The signal score weighs cross-source corroboration, recency, source weight and topic salience. How we rank stories.
Original reporting
This story summarises reporting published by Towards AI. Read the original article at Towards AI.