Prompt caching
Providers reuse the computed state (KV cache) for an identical prompt prefix, charging less for the cached part and lowering latency. The rule for all vendors: put static content first (tools, system prompt, reference documents), variable content last. Anything that changes the prefix (a timestamp, reordered tools) breaks the hit. See context-engineering.
Anthropic (Claude API)
- Automatic caching (top-level
cache_control) or explicit breakpoints on blocks; up to 4 explicit breakpoints per request. - TTL 5 minutes by default, 1 hour optional. Multipliers on base input price: 5-minute write 1.25x, 1-hour write 2x, read 0.1x (docs list 0.05x on Opus 5.5 and 0.025x on Fable/Mythos 5.1).
- Exception to flat long-context pricing: Claude Haiku 5.5 is priced by prompt length. A request whose prompt is over 100,000 tokens pays higher rates on every line, and the prompt length counts cache reads and writes, so a cache hit on a long prompt is still billed at the higher tier. Details and arithmetic: claude-haiku-5-5-pricing-and-caching.
- Minimum cacheable prompt: 512 tokens on newest models (Opus 5.5, Sonnet 5.5, Fable 5.1 etc.), 1,024 to 4,096 on older ones; below it the request still works, uncached.
- Invalidation cascades tools → system → messages. Changing tool definitions invalidates everything; changing
tool_choice, images oroutput_config.effortinvalidates messages; changing structured-output format also invalidates. - Usage fields:
cache_creation_input_tokens,cache_read_input_tokens,input_tokens.
OpenAI
- Automatic, prefix-based. For GPT-5.6 and later: 1,024-token minimum, retention
30m, cached input at 0.1x (0.05x on GPT-6.1 Sol) and a cache-write surcharge of 1.25x. Earlier models:in_memory(5-10 min idle) or24hretention, no additive write fee. prompt_cache_keyroutes or separates cache accounting; usage showscached_tokensandcache_write_tokens.
Google Gemini
- Implicit caching on by default for Gemini 2.5 and newer, plus explicit cache objects (not supported in the Interactions API). Minimums: 4,096 tokens for Gemini 3.5 to 3.8 Flash, 2,048 for 2.5 Flash/Pro. Docs say savings are passed on but the fetched page gave no percentage or storage price: check the pricing page.
Design implications
- Agents: the Manus team calls KV-cache hit rate “the single most important metric” for a production agent and avoids removing tools mid-run (mask instead).
- MCP servers should return tools in deterministic order to help hit rates (MCP 2026-07-28 spec).
- Pricing and minimums change by model; re-check before budgeting.
Sources
- https://platform.claude.com/docs/en/build-with-claude/prompt-caching (2026-09-30)
- https://developers.openai.com/api/docs/guides/prompt-caching (2026-09-30)
- https://ai.google.dev/gemini-api/docs/caching (2026-09-30)
- https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost (2026-09-30)
- https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus (2026-09-30)
- https://modelcontextprotocol.io/specification/2026-07-28/server/tools (2026-09-30)