Claude Haiku 5.5: Pricing, Migration Changes, and When to Use It

TL;DR
Claude Haiku 5.5 (claude-haiku-5-5) costs $0.10 input and $0.50 output per million tokens under 100K, with a 1M window and adaptive thinking. The pricing cliff, the tokenizer catch, and the ten migration steps from Haiku 4.5.
Last updated: October 7, 2026 - prices, benchmarks, and migration steps checked against Anthropic's announcement and the Claude Platform docs on launch day; community reaction read from the Hacker News launch thread.
Claude Haiku 5.5 (claude-haiku-5-5) is Anthropic's new small model, released October 7, 2026. It lists at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, with a 1M-token context window, 128K max output, and adaptive thinking. Use it for summaries, compaction, classification, routing, and subagent work, and keep Sonnet 5.5 or Opus 5.5 for hard agentic coding. The headline "around 75% cheaper than Haiku 4.5" is real but conditional, and the condition is a token-count cliff that the new tokenizer makes easier to hit than it looks.
Primary Documents#
| Source | Link |
|---|---|
| Anthropic announcement (October 7, 2026) | anthropic.com/claude-haiku-5-5 |
| Haiku 5.5 model overview (ids, limits, price rows) | platform.claude.com/.../haiku-5-5/overview |
| Haiku 5.5 migration guide | platform.claude.com/.../haiku-5-5/migration-guide |
| System card | anthropic.com/document/claude-haiku-5-5-system-card |
What Anthropic Says It Is For#
Anthropic positions Haiku 5.5 as "the cheapest, fastest, and most capable small model" it has shipped, aimed at "high-volume, cost-sensitive tasks": summaries, compactions, database queries, classification. It also pitches it as a subagent under Opus 5.5 or Sonnet 5.5, and for speed-sensitive work like live customer support and browser use. It is the first Haiku-class model with an adjustable effort setting, and default effort is medium.
Anthropic is explicit about the limit. Its own Terminal-Bench 4.0 chart shows Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding tasks", while Haiku 5.5 suits "more narrowly scoped tasks that might otherwise have been cost-prohibitive".
Launch-day extras, from the same announcement:
- Sonnet 5.5 cache reads drop 50%, from $0.20 to $0.10 per million tokens, which Anthropic says cuts Sonnet 5.5 cost on most agentic work by around 20%.
- Max and Team subscribers get a monthly Claude Platform API credit: $100 for Max 5x, $200 for Max 20x, up to $500 pooled for Team.
- The Python and TypeScript SDKs add beta computer use and browser use support.
Benchmarks (Vendor-Reported)#
All numbers below are from Anthropic's announcement table, not independent runs. GDPval-AA is an Artificial Analysis Elo score that Anthropic reports; treat the cross-vendor column as a vendor-chosen comparison.
| Benchmark | Haiku 5.5 | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 39.2% | 0.0% | 70.6% |
| FrontierCode 1.1 (main) | 46.4% | - | 52.1% (xhigh) |
| OSWorld 2.1 (offline subset) | 72.4% | 15.7% | 83.9% |
| Humanity's Last Exam (no tools) | 45.9% | 10.2% | 56.9% |
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1840 |
Read the Terminal-Bench row twice. A 39.2% against 70.6% is the gap that decides your routing rule: terminal-driven, multi-step coding stays on the bigger model, and Haiku takes the narrow, bounded steps.
Pricing: The 100K Cliff#
Per million tokens, from the model overview and announcement:
| Haiku 5.5 (up to 100K prompt) | Haiku 5.5 (over 100K) | Haiku 4.5 | Sonnet 5.5 | |
|---|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 | $2.00 |
| Output | $0.50 | $2.50 | $5.00 | $10.00 |
| Cache read | $0.01 | $0.05 | $0.10 | $0.10 (new) |
| 5-minute cache write | $0.125 | $0.625 | $1.25 | $2.50 |
Batch API takes 50% off input and output. Anthropic says 90% of Haiku 4.5 requests fell under 100K tokens, so for those the price is 90% lower; over 100K it is 50% lower. The "around 75% less on average" figure also bakes in the tokenizer change.
That tokenizer change is the catch. The migration guide says Haiku 5.5 uses the same tokenizer as Claude 4.7 and later, and "the same input text produces approximately 30% more tokens" than on Haiku 4.5. The price cutoff is counted in the new tokens. A prompt that was 90K tokens on Haiku 4.5 is roughly 117K on Haiku 5.5, so it lands in the upper price tier.
Worked example, list prices, assuming the documented ~30% inflation applies uniformly (it varies by content, so count your own prompts). One agent turn on Haiku 4.5 with a 100K-token prompt (90K cached reads, 10K fresh) and 4K output:
- Haiku 4.5: 90K cache reads = $0.009, 10K input = $0.010, 4K output = $0.020. Total $0.039.
- Haiku 5.5, same text (about 130K tokens, upper tier): 117K cache reads = $0.00585, 13K input = $0.0065, 5.2K output = $0.013. Total about $0.0254, roughly 35% cheaper.
Now a 60K-token turn on Haiku 4.5 (54K cached, 6K fresh, 4K output) costs $0.0314. On Haiku 5.5 it becomes about 78K tokens, still under the cliff: about $0.0041, roughly 87% cheaper. Same model, same discount headline, very different savings depending on whether you cross 100K. If you replay long conversations into a small model, prompt caching and trimming context to stay under the line matter more than the sticker price.
Migrating From Haiku 4.5: Ten Checklist Items#
The migration guide lists ten changes. The ones that will actually break code:
- Model id.
claude-haiku-5-5on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS;anthropic.claude-haiku-5-5on Bedrock. No date suffix and no alias. - Recount tokens. Re-measure prompts,
max_tokens, and cost estimates withmodelset toclaude-haiku-5-5. - Thinking.
thinking: {"type": "enabled", "budget_tokens": N}returns a 400. Use{"type": "adaptive"}and set effort throughoutput_config. Thinking tokens count towardmax_tokens, so a small limit can stop after a thinking block and before any text. - Select blocks by type, not position, because a response can start with a thinking block.
- Drop sampling parameters. Any
top_k, anytemperatureother than 1, anytop_pother than 0.99, or sending bothtemperatureandtop_p, returns a 400. - No assistant prefill. A final assistant turn is rejected with a 400 even with thinking off.
- Computer use (Claude API and Google Cloud) moves to the
computer_toolset_20260801toolset; the oldcomputer_20250124tool returns a 400. - Thinking blocks are account-bound. Replay them only through the account that produced them.
- Keep conversations append-only. Do not send thinking blocks back after editing
system,tools, or earlier messages. - Handle
stop_reason: "refusal". Haiku 5.5 runs safety classifiers and, per the guide, has no server-side fallback.
Priority Tier is not supported on Haiku 5.5. The request that replaces a Haiku 4.5 call looks like this, copied from the guide's before/after (not run by us, since no API key was available while writing):
{
"model": "claude-haiku-5-5",
"max_tokens": 16000,
"thinking": { "type": "adaptive" },
"output_config": { "effort": "medium" },
"messages": [{ "role": "user", "content": "..." }]
}
Anthropic's guidance for work that used to run without thinking is to pick a lower effort level, since at lower levels the model "can skip thinking entirely on simpler requests". For effort tradeoffs on the bigger models, see Fable 5 effort levels explained.
Where It Fits In An Agent Stack#
The practical routing rule, built from Anthropic's own positioning rather than from our testing:
- Haiku 5.5: compaction, summarization, classification, ticket triage, extraction, browser-use steps, and exploratory subagents that return a short report.
- Sonnet 5.5: the default coding model; the cache-read cut makes long loops cheaper. Our Sonnet 5.5 guide has the five breaking API changes (its cache-read row predates the cut).
- Opus 5.5: planning, review, and the hardest edits. See the Opus 5.5 guide.
If you orchestrate several agents, subagents vs agent teams vs workflows covers where a cheap worker model pays off. We did not benchmark Haiku 5.5 as a Claude Code subagent; Anthropic's Cognition quote says it works as a "sidekick" in Devin Fusion, which is a vendor-selected customer claim, not a measurement we can vouch for.
What People Are Actually Saying#
The Hacker News launch thread had 268 points and 124 comments an hour or so after the release, so these are first reactions.
- Cheap summarization is the use people see. One commenter says Anthropic has lacked a cost-effective model for summarization, compaction and RAG helpers, and that those workloads usually fit under 100K tokens. Another says it could displace GPT-6 Luna for speed-sensitive tasks.
- The 100K cutoff is contested. One commenter calls it "absurdly low" and easy to exceed with agents; a reply quotes Anthropic's claim that 90% of Haiku 4.5 traffic was under it. A third notes Claude's tokenizer counts more tokens than GPT's for the same text, which cuts both ways when comparing cutoffs. That matches the arithmetic above.
- Haiku 4.5 left a bad taste. One user called the old model a waste of time and money for delegated coding, and another uses Haiku only for first-pass bug triage before handing off to Sonnet. Haiku 5.5 has to earn back that trust on real tasks.
- Chinese open models remain the counter-case. One commenter runs GLM-5.3-Flash for almost everything on subsidized plan pricing; another replies that cheaper-per-token models often need many more tokens per task. Neither side posted numbers, so cost per finished task on your workload is the only fair test.
Also see Reddit's r/ClaudeAI launch thread, which was a single hour old when we looked and had little to summarize yet.
FAQ#
What is the Claude Haiku 5.5 model id?#
claude-haiku-5-5 on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS, and anthropic.claude-haiku-5-5 on Amazon Bedrock. It has no date suffix and no alias.
How much does Claude Haiku 5.5 cost?#
For prompts up to 100,000 tokens: $0.10 input, $0.50 output, $0.01 cache read, $0.125 five-minute cache write per million tokens. Over 100,000 tokens: $0.50, $2.50, $0.05 and $0.625. The Batch API is 50% off.
Is Haiku 5.5 better than Sonnet 5.5 for coding?#
No. Anthropic reports 39.2% on Terminal-Bench 4.0 for Haiku 5.5 against 70.6% for Sonnet 5.5, and tells developers to keep Sonnet and Opus for complex agentic coding. Use Haiku for bounded steps and subagents.
Why is Haiku 5.5 not 75% cheaper for my workload?#
The new tokenizer produces about 30% more tokens for the same text, and the lower price tier ends at 100K tokens. Prompts that grow past that line pay the over-100K rates. Count tokens with claude-haiku-5-5 before projecting savings.
What breaks when I switch from Haiku 4.5?#
Fixed thinking budgets, non-default sampling parameters, assistant prefill, and the old computer use tool all return 400 errors. Switch to adaptive thinking, drop sampling parameters, end messages with a user turn, and handle refusal stop reasons.
Continue Reading#
- Claude Sonnet 5.5 Developer Guide - the mid-tier sibling and its five API changes
- Claude Opus 5.5 Developer Guide - the top of the 5.5 family
- Claude Haiku 4.5 - the model this one replaces
- Prompt Caching for the Claude API - cache math that decides your real bill
- Claude Batch API in Production - the 50% discount path for offline work
Sources#
| Source | What it supports |
|---|---|
| Anthropic, "Introducing Claude Haiku 5.5" (October 7, 2026) | Positioning, vendor benchmarks, pricing table, Sonnet 5.5 cache-read cut, API credits, SDK updates |
| Claude Haiku 5.5 model overview | Model ids, 1M context, 128K output, price tiers, Batch discount, default effort |
| Claude Haiku 5.5 migration guide | The ten migration items, tokenizer inflation, 400 errors, refusal behavior, Priority Tier |
| Hacker News launch thread | Community reaction quoted above (read October 7, 2026) |
Get the next comparison like this in your inbox
One email a week on News and the rest of the AI dev stack. Free.
Read next on Claude Code
Claude Sonnet 5.5 Developer Guide: Pricing, Benchmarks, and the Five API Changes
Claude Sonnet 5.5 (claude-sonnet-5-5) is Anthropic's new mid-tier model: $2/$10 per million tokens, 70.6% on Terminal-Bench 4.0, 1M context, now GA in GitHub Copilot and on Vercel AI Gateway. The pricing math, the five breaking API changes, and where it fits next to Opus 5.5 and Sonnet 5.
11 min readClaude Opus 5.5 Developer Guide: API Examples, Claude Code Setup, Pricing, and When to Use It
Claude Opus 5.5 (claude-opus-5-5) is Anthropic's new default Opus: $4/$20 per million tokens, $0.20 cache reads, 1M context, thinking always on with medium default effort. Runnable TypeScript and Python SDK examples, Claude Code setup, before/after prompts, and a decision guide vs Sonnet 5, Haiku 4.5, and Fable 5.1.
11 min readClaude Haiku 4.5: Near-Frontier Intelligence at a Fraction of the Cost
Anthropic's Claude Haiku 4.5 delivers Sonnet 4-level coding performance at one-third the cost and twice the speed. Here is what developers need to know.
5 min readTechnical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








