Claude Haiku 5.5

Anthropic’s fastest and cheapest current Claude model, released 2026-10-07. It succeeds Claude Haiku 4.5 and is built for “high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks” (Anthropic docs). Anthropic’s announcement calls it “the cheapest, fastest, and most capable small model we’ve ever released” (vendor claim). Family hub: anthropic-claude; maker: Anthropic.

Key facts (Anthropic docs, read 2026-10-09)

Model IDclaude-haiku-5-5 on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS; anthropic.claude-haiku-5-5 on Amazon Bedrock. A fixed ID with no date suffix and no separate alias
Released2026-10-07 (model page and release notes)
Context window / max output1M tokens / 128K tokens (up from 200K / 64K on Haiku 4.5); up to 300K output on the Batch API with the output-300k-2026-03-24 beta header
ThinkingAdaptive thinking, on by default; effort parameter with default medium; it is the first Haiku with effort levels (low, medium, high, xhigh, max)
Input / outputText and images in, text out
Knowledge cutoffJune 2026 (reliable and training data)
LatencyFastest in the current lineup (Anthropic’s comparative rating)
RetirementNot sooner than 2027-10-07 on Anthropic-operated platforms
Not supportedPriority Tier; server-side refusal fallback

Pricing structure: the 100K-token threshold

Unlike the other current Claude models, which include the full 1M-token window at one rate, Haiku 5.5 is priced by prompt length: a request whose prompt is over 100,000 tokens pays a higher rate on every line (input, output, cache writes and cache reads, and the Batch API). The prompt length counts all input tokens, including cache reads and cache writes, so caching does not avoid the higher tier. Rates, worked examples and comparisons are in claude-haiku-5-5-pricing-and-caching.

What is new compared with Haiku 4.5

  • Adaptive thinking and effort replace the manual thinking budget.
  • 1M context and 128K output.
  • Browser use tool (browser_toolset_20260801) and, on the Claude API and Google Cloud, computer use through the computer_toolset_20260801 toolset only.
  • Safety classifiers can decline requests (stop_reason: "refusal" with a category of cyber, frontier_llm, bio or general_harms); there is no server-side fallback, and an identical retry usually refuses again.
  • A newer tokenizer: the same text counts as about 30% more tokens than on Haiku 4.5, so token budgets, max_tokens and cost estimates need recounting.

Breaking changes when migrating from Haiku 4.5 (migration guide)

  1. thinking: {"type": "enabled", "budget_tokens": N} returns a 400 error; use adaptive thinking and effort.
  2. Non-default temperature, top_p and top_k return a 400 error (if present, temperature must be 1 and top_p 0.99).
  3. Assistant message prefill returns a 400 error; end messages with a user turn.
  4. computer_20250124 returns a 400 error; use the toolset.
  5. A thinking block sent back after a change to system, tools or earlier messages returns a 400 error, so keep conversations append-only. Thinking blocks also work only in the account that produced them (or a linked account).
  6. Responses can start with thinking blocks and thinking text is omitted by default (set thinking.display to "summarized" to see it), so select content blocks by type.

Vendor benchmark claims (Anthropic announcement, 2026-10-07)

The announcement compares Haiku 5.5 with Haiku 4.5, GPT-6 Luna and Sonnet 5.5; the column order below was inferred from the flattened page text and methodology sits in the system card, so treat the figures as vendor-reported:

BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
GDPval-AA v2.1 (Elo)162073514371840
OSWorld 2.1 (offline subset)72.4%15.7%48.9%83.9%
Humanity’s Last Exam (no tools)45.9%10.2%not reported56.9%
Terminal-Bench 4.039.2%0.0%16.4%70.6%

Anthropic says Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding. Customer figures in the announcement (for example Asana’s reported latency reduction, Box’s score gain) are the customers’ own, reported by Anthropic.

Prompting notes (Anthropic’s Haiku 5.5 guide)

  • Effort: medium is the default; low suits chat and simple high-volume work; high suits instruction-following and long agent tasks; at xhigh and max, thinking cannot be turned off.
  • Caching interaction: changing the top-level effort between requests invalidates the prompt cache for the conversation’s messages; a per-message effort change (beta) keeps it.
  • Search: give the model today’s date when it has a search tool; it sometimes needs a nudge to search at low effort.
  • Agents: at low effort in long agent prompts it can stop early or report code changes without running a check; the guide supplies prompt text for both.
  • Mid-turn user input: deliver it as a user turn, not inside a tool_result, because the model resists prompt injection through tool results.

Fit (opinion)

Strong candidate for routing, classification, extraction, summarisation and subagent work, where its rate and speed matter and Anthropic’s own table shows large gains over Haiku 4.5. Keep requests under 100,000 tokens where you can; above that, compare against Sonnet 5.5 on your own evals, since Sonnet’s rates do not step up with prompt length.

Open items

  • The benchmark column order is inferred; no independent evaluations were found.
  • How the other cloud platforms (Bedrock, Google Cloud, Foundry) price the over-100K tier is not confirmed here; the pricing page lists them separately.
  • The announcement does not state an AI Safety Level and does not explain why 100,000 tokens was chosen as the threshold.

Sources (read 2026-10-09)