Claude Haiku 5.5 pricing and prompt caching
Pricing snapshot, 2026-10-09
The amounts below are from Anthropic’s pricing page and Haiku 5.5 model page as read on 2026-10-09 (USD per million tokens, MTok). This note’s subject is pricing, so it keeps the amounts; check https://platform.claude.com/docs/en/about-claude/pricing before relying on them. Model overview: claude-haiku-5-5.
What is special: a 100,000-token step
Most current Claude models include the full 1M-token context at one rate. Haiku 5.5 does not. Anthropic’s pricing page: “Claude Haiku 5.5 is priced by prompt length: a request whose prompt is over 100,000 tokens pays higher prices. A request’s prompt length counts all of its input tokens, including cache reads and cache writes. Each request is priced on its own: a request over the threshold pays the higher prices even when part of its prompt is a cache hit, and earlier requests keep the prices they were billed at.”
So there is a price cliff at 100,000 tokens, it applies to the whole request, and prompt caching does not get you under it, because cached tokens count toward the length.
Rates
| Line (USD / MTok) | Prompt up to 100,000 tokens | Prompt over 100,000 tokens | Step |
|---|---|---|---|
| Base input | 0.10 | 0.50 | 5x |
| 5-minute cache write | 0.125 | 0.625 | 5x |
| 1-hour cache write | 0.20 | 1.00 | 5x |
| Cache read (hit or refresh) | 0.01 | 0.05 | 5x |
| Output | 0.50 | 2.50 | 5x |
| Batch input / output | 0.05 / 0.25 | 0.25 / 1.25 | 5x |
Within each tier the usual caching ratios hold: a 5-minute write is 1.25x base input, a 1-hour write 2x, and a read 0.1x. Batch (a 50% discount) and data-residency multipliers stack on top, per the pricing page. Output tokens are priced by the prompt’s tier, not by the output length.
What counts toward the 100,000
- All input tokens of the request: new input, cache reads and cache writes (
input_tokens + cache_read_input_tokens + cache_creation_input_tokens; the caching docs define total input this way). - “Up to 100,000” is the lower tier and “over 100,000” the higher, so a prompt of exactly 100,000 tokens is billed at the lower rate (reading of the wording).
- Tokens are counted with Haiku 5.5’s tokenizer, which gives about 30% more tokens for the same text than Haiku 4.5, so the threshold equals roughly 55,000 words of English (Anthropic’s conversion for the current tokenizer is about 555,000 words per 1M tokens).
Worked examples (arithmetic from the rates above)
| Request | Calculation | Cost |
|---|---|---|
| 100,000 uncached input tokens | 100,000 x 0.10 / 1,000,000 | $0.0100 |
| 100,001 uncached input tokens | 100,001 x 0.50 / 1,000,000 | $0.0500 (5x for one more token) |
| 100,000 tokens, all cache reads | 100,000 x 0.01 / 1,000,000 | $0.0010 |
| 100,001 tokens, all cache reads | 100,001 x 0.05 / 1,000,000 | $0.0050 |
| 90K prompt (80K cache read + 10K new) | 80,000 x 0.01 + 10,000 x 0.10 | $0.0018 |
| 120K prompt (110K cache read + 10K new) | 110,000 x 0.05 + 10,000 x 0.50 | $0.0105 (5.8x the cost for 1.33x the tokens) |
| Writing a 90K prefix to the 5-minute cache | 90,000 x 0.125 / 1,000,000 | $0.01125 |
| Writing a 150K prefix to the 5-minute cache | 150,000 x 0.625 / 1,000,000 | $0.09375 |
| 2,000 output tokens | tier-1 vs tier-2 | 0.0050 |
A cache read on a long prompt (0.05) is still cheaper than uncached input on a short one (0.10), but a request that crosses the line also pays 5x on its output and on any new tokens.
How it compares
| Model (USD / MTok) | Input | Output | Cache read | 5-min write | Long-context step |
|---|---|---|---|---|---|
| Haiku 5.5, up to 100K | 0.10 | 0.50 | 0.01 | 0.125 | prompt-length tier |
| Haiku 5.5, over 100K | 0.50 | 2.50 | 0.05 | 0.625 | |
| Haiku 4.5 | 1.00 | 5.00 | 0.10 | 1.25 | none (200K window) |
| Sonnet 5.5 | 2.00 | 10.00 | 0.10 | 2.50 | none across 1M |
| Opus 5.5 | 4.00 | 20.00 | 0.20 | 5.00 | none across 1M |
- Even above 100K, Haiku 5.5 is cheaper per token than Sonnet 5.5 on every line (4x on input, output and writes; 2x on cache reads), and 50% cheaper than Haiku 4.5. Anthropic states the same: Haiku 5.5 costs 90% less than Haiku 4.5 up to 100K and 50% less above, an average reduction of “around 75%”, and about 90% of Haiku 4.5 requests were within 100K (vendor figures from the announcement).
- Tokenizer effect (my arithmetic): because the same text is about 30% more tokens, the saving on identical text is about 87% for short prompts (1.3 x 0.10 against 1.00) and about 35% above 100K (1.3 x 0.50 against 1.00).
- Cheaper does not mean equivalent: on Anthropic’s own table Sonnet 5.5 and Opus 5.5 are stronger on complex agentic work.
Other caching details that affect cost
- Minimum cacheable prompt on Haiku 5.5: 512 tokens; shorter prompts are processed uncached without an error.
- Changing the top-level
effortbetween requests invalidates the cached conversation messages; use the per-message effort change (beta) to keep the cache. - Thinking blocks stay valid only while earlier turns are unchanged, so keep conversations append-only, which also helps cache hits.
- Priority Tier is not supported on Haiku 5.5.
Practical guidance (opinion)
- Watch the
usagetotals: if input plus cache reads plus cache writes will pass 100K, expect 5x rates on the entire request including output. - Keep long-running agents under the line with compaction, context editing or retrieval instead of carrying the full history (prompt-caching, context-compaction-and-memory).
- If your prompts routinely exceed 100K, price the same workload on Sonnet 5.5, whose rates do not step up, and compare quality as well as cost.
- Split bulk document jobs into requests that each stay under 100K where the task allows it.
- For short, repeated prompts the lower tier is very cheap: the cache read rate is one two-hundredth of the price of a Sonnet 5.5 input token (0.01 against 2.00).
Open items
- Whether Amazon Bedrock, Google Cloud and Microsoft Foundry apply the same 100K step is not confirmed; the pricing page lists partner platforms separately, and Claude Platform on AWS bills at Claude API rates.
- The caching documentation does not explain how the threshold interacts with cache breakpoints; the pricing page’s rule above is the only statement found.
- Anthropic gives no reason for choosing 100,000 tokens.
- How thinking tokens enter the threshold is not stated (the page speaks only of prompt length).
Sources (read 2026-10-09)
- https://platform.claude.com/docs/en/about-claude/pricing (model pricing, prompt caching, long context, batch tables)
- https://platform.claude.com/docs/en/models/haiku-5-5/overview
- https://platform.claude.com/docs/en/build-with-claude/prompt-caching (minimum cacheable length, usage fields)
- https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-haiku-5-5 (effort and cache)
- https://www.anthropic.com/claude-haiku-5-5 (announcement; vendor figures)