Anthropic and OpenAI offer 90% AI prompt caching discounts for agents
Anthropic, OpenAI, and Google have introduced prompt and KV caching layers that offer up to 90% cost reductions for repeated input prefixes in AI agent workflows. These optimizations allow developers to significantly reduce latency and operational costs by reusing computed attention vectors for stable system prompts and tool definitions.
Key Takeaways
- Anthropic offers explicit cache control with 90% read discounts and a 25% premium on cache writes.
- OpenAI GPT-5.6 and newer models automatically cache eligible prefixes with a 30-minute default time-to-live.
- Google Gemini 2.5 provides implicit caching by default alongside explicit named caches for granular control.
- Project Discovery increased cache hit rates from 7% to 74% by moving dynamic content to the end of user messages.
Why It Matters
The shift toward deep discounts for cached tokens fundamentally changes the unit economics for streaming platforms deploying AI-driven personalization and automated metadata tagging. By rewarding stable prefixes, these pricing models incentivize developers to standardize system prompts and tool definitions, effectively turning GPU memory into a strategic asset. Within the broader ecosystem, this optimization bridges the gap between high-latency RAG systems and real-time user interactions, allowing for more complex agentic behavior without linear cost increases. Watch for whether open-source engines like VLLM and Moon Cake can maintain their throughput advantages as proprietary providers move toward fully implicit, infrastructure-level caching.
Additional Context
Anthropic's prompt caching feature, launched in late 2024, has become a key cost lever for high-volume AI workloads. In June 2026, Ericsson launched its AI in RAN commercial software subscription claiming up to 20% higher downlink throughput across more than 15 live deployments, demonstrating how agentic AI systems with stable system prompts benefit from caching economics at scale. For streaming platforms running continuous personalization pipelines, the same prefix-reuse pattern applies: stable tool definitions and system instructions can be cached across millions of inference calls, reducing per-request costs by an order of magnitude.
The competitive dynamics around prompt caching extend beyond Anthropic and OpenAI. Nokia and Google Cloud announced Gemini-powered AI agents for telco network troubleshooting at DTW IGNITE 2026, with six specialized agents targeting alarms, KPIs, anomaly detection, and remediation. These agents rely on fixed system prompts and tool schemas that are prime candidates for prompt caching, illustrating how the discount model directly subsidizes agentic architectures in production. Google's Gemini platform, which also offers context caching, competes directly with Anthropic and OpenAI on this pricing dimension, and the Nokia deployment signals enterprise demand for cost-efficient agent orchestration.
On the infrastructure side, the divergence between proprietary and open-source approaches to caching is sharpening. Ericsson's strategy of running only forward error correction on GPUs while keeping other Layer 1 functions on CPUs contrasts with Nokia's full GPU-based approach, mirroring the broader debate over where caching and inference compute should reside. For streaming AI workloads, open-source engines like VLLM and Moon Cake implement their own KV-cache management strategies that can outperform proprietary APIs on throughput for specific batch sizes, but they lack the implicit, infrastructure-level caching that Anthropic and OpenAI now offer at discounted rates. Ericsson's CTO Erik Ekudden highlighted that uplink traffic could triple over the next five years driven by AI glasses, sensors, and real-time video, underscoring that the volume of AI inference calls requiring cached prefixes will grow substantially as edge and streaming use cases expand.
Read full article at m.youtube.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source