Prompt caching

Providers reuse the computed state (KV cache) for an identical prompt prefix, charging less for the cached part and lowering latency. The rule for all vendors: put static content first (tools, system prompt, reference documents), variable content last. Anything that changes the prefix (a timestamp, reordered tools) breaks the hit. See context-engineering.

Anthropic (Claude API)

  • Automatic caching (top-level cache_control) or explicit breakpoints on blocks; up to 4 explicit breakpoints per request.
  • TTL 5 minutes by default, 1 hour optional. Multipliers on base input price: 5-minute write 1.25x, 1-hour write 2x, read 0.1x (docs list 0.05x on Opus 5.5 and 0.025x on Fable/Mythos 5.1).
  • Exception to flat long-context pricing: Claude Haiku 5.5 is priced by prompt length. A request whose prompt is over 100,000 tokens pays higher rates on every line, and the prompt length counts cache reads and writes, so a cache hit on a long prompt is still billed at the higher tier. Details and arithmetic: claude-haiku-5-5-pricing-and-caching.
  • Minimum cacheable prompt: 512 tokens on newest models (Opus 5.5, Sonnet 5.5, Fable 5.1 etc.), 1,024 to 4,096 on older ones; below it the request still works, uncached.
  • Invalidation cascades tools → system → messages. Changing tool definitions invalidates everything; changing tool_choice, images or output_config.effort invalidates messages; changing structured-output format also invalidates.
  • Usage fields: cache_creation_input_tokens, cache_read_input_tokens, input_tokens.

OpenAI

  • Automatic, prefix-based. For GPT-5.6 and later: 1,024-token minimum, retention 30m, cached input at 0.1x (0.05x on GPT-6.1 Sol) and a cache-write surcharge of 1.25x. Earlier models: in_memory (5-10 min idle) or 24h retention, no additive write fee.
  • prompt_cache_key routes or separates cache accounting; usage shows cached_tokens and cache_write_tokens.

Google Gemini

  • Implicit caching on by default for Gemini 2.5 and newer, plus explicit cache objects (not supported in the Interactions API). Minimums: 4,096 tokens for Gemini 3.5 to 3.8 Flash, 2,048 for 2.5 Flash/Pro. Docs say savings are passed on but the fetched page gave no percentage or storage price: check the pricing page.

Design implications

  • Agents: the Manus team calls KV-cache hit rate “the single most important metric” for a production agent and avoids removing tools mid-run (mask instead).
  • MCP servers should return tools in deterministic order to help hit rates (MCP 2026-07-28 spec).
  • Pricing and minimums change by model; re-check before budgeting.

Sources