Cheapest Subagent Model: Haiku 5.5 or GPT-6 Luna? Routing, Compaction and Cost per Task

TL;DR
Which cheap model belongs under your orchestrator? Haiku 5.5 vs GPT-6 Luna as a subagent or worker: routing by prompt length, compaction, tool-call gotchas and cost per agent task.
Last updated: October 9, 2026 - new page on choosing between the two as a subagent or worker model. Prices, limits and availability checked against Anthropic's Haiku 5.5 announcement and model overview, OpenAI's GPT-6 Luna model page and Codex models doc, the GitHub and Vercel changelogs, the Claude Code model-configuration docs, and Artificial Analysis's Haiku 5.5 write-up on this date.
This page answers one question: which of the two cheap models should run under your orchestrator as a subagent or worker in an agent harness such as Claude Code, Codex, Copilot or your own loop. For the straight list-price comparison, and where both sit against DeepSeek V4.1-Flash and Gemini 3.5 Flash, see Budget AI Coding Models, October 2026.
The short answer: use Claude Haiku 5.5 for short, bounded subagent calls that stay under 100,000 tokens per request, and use GPT-6 Luna for long-running workers that replay a growing history. Both list at $0.10 per million input tokens and $0.50 per million output, with $0.01 cache reads, so the sticker settles nothing. For a worker, what settles it is how long each call's prompt gets, how many tokens each model spends to finish the task, and how cleanly it fits the harness's tool-calling path.
Under 100K, a cached Haiku 5.5 subagent costs the same as Luna at equal token counts and, on both Anthropic's and Artificial Analysis's numbers, scores higher on agentic coding. Past 100K, Haiku bills 5x on every line of the request while Luna holds its base price to 272K. In our worked example below, a 30-turn worker that grows to 194K tokens costs 3.5x more on Haiku than on Luna before you count the tokenizer, and 5-7x more after. So the routing rule is prompt length first, model brand second, and the harness setting that matters most is when it compacts.
For the Haiku 5.5 migration checklist and the full Anthropic pricing table, see the Haiku 5.5 release guide. For routing between Haiku and its larger siblings, see Opus 5.5 vs Sonnet 5.5 vs Haiku 5.5.
The Two Price Lines a Worker Crosses#
A worker's bill is mostly cache reads and output tokens, turn after turn, so only a few rows of each price sheet matter. Per million tokens, standard tier, from the Haiku 5.5 model overview and the GPT-6 Luna model page, checked October 9, 2026. The full side-by-side sheet, Batch rates included, is in the budget models roundup.
| Row a worker pays every turn | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Base price holds until | 100K-token prompt | 272K-token request |
| Cache read, below / above the line | $0.01 / $0.05 | $0.01 / $0.02 |
| Cache write, below / above | $0.125 / $0.625 (5-minute) | $0.125 / $0.25 |
| Output, below / above | $0.50 / $2.50 | $0.50 / $0.75 |
| Default effort (API) | medium | medium |
| Context / max output | 1M / 128K | 1,050,000 / 128,000 |
OpenAI's page says Luna's surcharge applies "for the full request"; Anthropic's overview keys Haiku's rate to prompt length, and the Reddit numbers we test below only reconcile if the whole request bills at the upper rate, so that is how we model it. The practical point for a harness: Luna's $0.01 cache read matches Haiku only below 100K. Above it, a Haiku cache read costs $0.05, and a long agent loop is mostly cache reads.
Three Things That Set Cost per Agent Task#
Three levers turn the same list price into different bills for the same agent task.
The 100K cliff. Every turn of an agent loop resends the conversation so far. Once that prompt passes 100,000 tokens, every remaining Haiku turn pays $0.05 per cached token instead of $0.01, and $2.50 per output token instead of $0.50. Luna's equivalent line is 272K, so a loop that peaks at 200K never leaves Luna's base tier.
The tokenizer. Anthropic's overview says Haiku 5.5 uses the newer Claude tokenizer, so "the same text counts as approximately 30% more tokens" than on Haiku 4.5. That comparison is against Haiku 4.5, not against OpenAI. Nobody has published a vendor number for Haiku 5.5 versus Luna on the same text. One Hacker News commenter estimates Claude's 100K tokens equal about 60-65K GPT tokens, which would mean Claude counts around 1.6x more. Treat that as a lead to test with Anthropic's token-counting endpoint on your own prompts, not as a fact. Either way, the practical effect is that Haiku's 100K line arrives earlier in real text than the number suggests.
Tokens spent per task. Thinking tokens bill as output. Artificial Analysis measured Haiku 5.5 at max effort using about 162K output tokens per Intelligence Index task, roughly 3x GPT-6 Luna at max (about 50K). At matched capability the gap narrows: Haiku at high scored 38 with about 55K tokens per task, the same score Luna reached at max with about 50K. Effort labels do not map across vendors, so compare at equal quality, not at equal setting names.
Worked Example: One Subagent, Two Shapes#
Our arithmetic at list prices, with prompt caching on: each turn reads the previous prompt from cache and writes only the new tool results, at the 1.25x cache-write rate both vendors charge. We treat the previous turn's output as part of the cached prefix; strictly it is a cache write, which adds about 5% to the cached rows and leaves the ratios almost unchanged. "Same count" assumes both models count the text identically. The +30% and +60% columns inflate only the Haiku side, as a sensitivity range for the tokenizer question above. Every number is ours, computed from the two price sheets; none is a measurement.
Shape A, a focused subagent. A 12K-token starting prompt (system prompt, tool definitions, task brief), 15 turns, each adding 3K tokens of tool results and 600 output tokens. The final prompt is 62.4K tokens.
Shape B, a long-running worker. A 20K-token starting prompt, 30 turns, each adding 5K tokens of tool output and 1K output tokens. The final prompt is 194K tokens.
| Shape | GPT-6 Luna | Haiku 5.5, same count | Haiku 5.5, +30% | Haiku 5.5, +60% |
|---|---|---|---|---|
| A: focused subagent, cached | $0.0163 | $0.0163 (1.0x) | $0.0212 (1.3x) | $0.0261 (1.6x) |
| A: focused subagent, no cache | $0.0603 | $0.0603 (1.0x) | $0.0784 (1.3x) | $0.0965 (1.6x) |
| B: long worker, cached | $0.0661 | $0.2302 (3.5x) | $0.3402 (5.1x) | $0.4415 (6.7x) |
| B: long worker, no cache | $0.3360 | $1.3216 (3.9x) | $1.9136 (5.7x) | $2.4525 (7.3x) |
Here is one turn by hand so you can check the method. The last turn of Shape A, at equal counts, is identical on both models: 59,400 cached tokens x $0.01 = $0.000594, plus 3,000 new tokens written to cache x $0.125 = $0.000375, plus 600 output x $0.50 = $0.0003. Total $0.00127.
The last turn of Shape B is where they split. The prompt is 194,000 tokens, 189,000 of them cached. On Luna, still under 272K: 189,000 x $0.01 + 5,000 x $0.125 + 1,000 x $0.50 = $0.00302. On Haiku, over 100K: 189,000 x $0.05 + 5,000 x $0.625 + 1,000 x $2.50 = $0.01508. Exactly 5x, and 16 of Shape B's 30 turns sit past the line at equal counts (20 at +30%, 22 at +60%).
Three readings fall out of the table:
- Under the line, it is a capability decision. Shape A costs the same on both at equal counts, and at most 1.6x more on Haiku under the steepest tokenizer assumption, which still peaks just under 100K. At these sizes, pay for whichever model needs fewer retries.
- Over the line, it is a context-management decision. Shape B on Haiku costs 3.5x to 7.3x Luna. You can claw most of that back by compacting or restarting the worker before 100K, which is the job prompt caching and context trimming were already doing.
- Caching narrows the gap at the bottom but not across the cliff. Without caching, both shapes cost about 4-6x more in absolute terms, and the Haiku multiple on Shape B gets worse, because uncached input pays the 5x on the full prompt every turn.
Testing the "12x More" Reddit Claim#
The community signal behind this page is an r/ClaudeAI post titled Claude Haiku 5.5 cost 12x more than GPT-6 Luna for the same voxel pagoda, posted on launch day. The same comparison went out on X from atomic.chat, an app that runs both models through their APIs, and drew more than 2,500 likes. It is a single creative build, not an agent benchmark, and it comes from a company selling a multi-model app, but it is a public, checkable example of a long agent session crossing Haiku's price line. We could not open the full Reddit thread; the figures below come from the post's results table and the atomic.chat post. Haiku 5.5 ran at xhigh effort: 268.8M input tokens of which 262.7M were cache reads, 4.46M output tokens, $24.96. Luna cost $1.91 for the same scene. Neither source gives a confirmed token count for the Luna run, so we leave it out. Treat all of these as the poster's numbers, not ours.
Then check whether the Haiku row is even possible at list prices:
- If every request had stayed under 100K: 262.7M x $0.01 + 6.1M x $0.10 + 4.46M x $0.50 = $5.47.
- If every request had been over 100K: 262.7M x $0.05 + 6.1M x $0.50 + 4.46M x $2.50 = $27.34.
- Reported: $24.96. That sits about 89% of the way from the floor to the ceiling, which is what you would expect from an agent session that crossed 100K early and spent most of its turns above it. Counting the uncached tokens at the cache-write rate instead moves that to about 86%.
So the row reconciles, and it decomposes cleanly. The same tokens kept under 100K would have cost about $5.47, roughly 2.9x Luna's $1.91. The cliff multiplied that by about 4.6x to reach $24.96. 2.9 x 4.6 is about 13, which matches $24.96 / $1.91 (13.1x); the posts round it to 12x.
What that means:
- The claim is arithmetically sound. Nothing about it requires a billing bug or a hidden fee.
- It is not a general rule. The 4.6x comes from the session running past 100K, and the run was at
xhigheffort rather than Haiku'smediumdefault. A bounded subagent that stays under the line loses the cliff multiplier entirely, leaving only the token-volume gap. - The token-volume gap is real too. Under 100K the two models charge identical rates, so the remaining 2.9x can only come from Haiku spending more cost-weighted tokens on the job, assuming Luna's run stayed under its own 272K line. That is in line with Artificial Analysis finding roughly 3x more output tokens per task at
max. Whether Haiku's pagoda was better is a separate question the cost table cannot answer. - For a harness, the lesson is compaction. A worker that compacts or restarts before 100K removes the 4.6x. Only the token-volume gap is left, and that one you test against your own tasks.
Worker Benchmarks: Vendor-Reported and Independent#
Anthropic's Haiku 5.5 announcement includes a GPT-6 Luna column. These are Anthropic-reported numbers about a competitor's model; OpenAI's Luna page publishes no benchmark scores to check them against. The rows that bear on worker jobs are agentic coding and computer use; the knowledge-work and visual rows are in the budget models roundup.
| Benchmark (Anthropic-reported) | Claude Haiku 5.5 | GPT-6 Luna | Sonnet 5.5, for reference |
|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 39.2% | 16.4% | 70.6% |
| FrontierCode 1.1, Main (agentic coding) | 46.4% | 42.4% | 52.1% (xhigh) |
| OSWorld 2.1, offline subset (computer use) | 72.4% | 48.9% | 83.9% |
Artificial Analysis ran its own evaluations, which are the closest thing to an independent check:
- Intelligence Index: Haiku 5.5 at
maxscores 43, against 38 for GPT-6 Luna. - Terminal-Bench 4.0: 33% for Haiku 5.5 against 13% for Luna. Lower than Anthropic's numbers for both, but the same direction and a similar ratio.
- Factual knowledge: Luna is ahead on AA-Omniscience accuracy (44% vs 36%), while Haiku's hallucination rate is much lower (40% vs 77%).
- AutomationBench-AA: Haiku scores 35% against 53-60% for Luna and other small models. Artificial Analysis says a pre-release over-refusal issue likely understates Haiku and plans to re-run it.
- Cost caveat: at the time of writing, Artificial Analysis said its site "does not yet reflect tiered pricing", so its provisional Haiku 5.5 cost figures leave out the over-100K step.
The pattern is consistent: Haiku 5.5 is the stronger model on agentic coding and computer use, by a wide margin on terminal work, and Luna is the more token-efficient one. Neither is in Sonnet 5.5's league for hard coding, and Anthropic says so itself: Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding tasks".
Tool Calls and Harness Fit#
Neither vendor publishes a tool-call error rate for these two models, so treat the following as the known gotchas, not a reliability score.
- Luna's tool calls belong on the Responses API. The model page says Chat Completions supports function calling with Luna only when
reasoning_effortisnone. A harness that drives tools through Chat Completions with reasoning on will not get tool calls from Luna. - Haiku 5.5 rejects sampling parameters. It returns a 400 for any non-default
temperature,top_portop_k, so a harness that hardcodes a temperature for its workers breaks on the switch from Haiku 4.5. The release guide lists the other breaking changes. - The one workflow benchmark where Luna leads has an asterisk. On Artificial Analysis's AutomationBench-AA, Haiku 5.5 scores 35% against 53-60% for Luna and other small models. Artificial Analysis says a pre-release safety issue made Haiku over-refuse, expects the score to rise, and plans to re-run it. Until then, test refusals on your own tool workflows.
- Hallucination cuts the other way. On AA-Omniscience, Haiku's hallucination rate is 40% against Luna's 77%, which matters for a worker whose output feeds the next step without a human check.
- Wire format. Under a Claude orchestrator, Haiku speaks the same API and tool schema as the lead model, so routing a subagent to it is a model-id change. Under an OpenAI orchestrator, the same is true of Luna.
Where Each One Runs This Week#
Both launched into the tools developers already use, which matters more for a worker model than a benchmark row.
Claude Haiku 5.5
- Claude API and clouds:
claude-haiku-5-5on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS;anthropic.claude-haiku-5-5on Amazon Bedrock (overview). - GitHub Copilot: generally available, rolling out gradually, for Copilot Pro, Pro+, Max, Business and Enterprise. GitHub positions it for "subagents, quick edits, and terminal tasks" and bills it at provider list pricing under usage-based billing (changelog). Our Copilot billing guide covers what that means for your seat.
- Vercel AI Gateway: available as
anthropic/claude-haiku-5.5, at provider pricing with no markup (changelog). - Claude Code: the
haikualias resolves to Haiku 5.5 on the Anthropic API, with Claude Code v2.1.293 or later required (model configuration). The same page warns that Haiku 5.5 sessions auto-compact at about 967K tokens by default, far past the 100K price line.
GPT-6 Luna
- OpenAI API:
gpt-6-lunaon the Responses, Chat Completions and Batch endpoints, among others (model page), with the function-calling caveat above. - Codex: OpenAI's models doc recommends Luna "for focused, repeatable tasks" and GPT-6.1 Sol for complex coding. On Free and Go plans, Luna is the replacement OpenAI names for GPT-5.5, which retires from Codex on October 14, and it is the named replacement for the retired
gpt-5.4-mini. OpenAI suggests starting Luna athigheffort; it supports up tomax, but not Ultra. Our Codex usage limits guide shows how many more Luna messages fit in a five-hour window than Sol or Astra.
The asymmetry: Luna is a recommended model inside OpenAI's own coding agent, and the one Free and Go users get, while Haiku 5.5 is a model you opt into in Claude Code, Copilot or a gateway. If you live in Codex, Luna is already in your loop. If you live in Claude Code or Copilot, Haiku is the one-line switch.
Run It Monday#
Commands below are from the official docs; we have not run them.
Claude Code, per the model configuration page. Start a session on Haiku 5.5, then set the auto-compact window to 100K, the documented minimum (the docs' own example is 500k), so the session compacts near the price line instead of at about 967K:
claude --model claude-haiku-5-5
/autocompact 100k
The docs say /autocompact saves the window per model under modelSettings, and accepts 100K to 1M. To make Haiku the default for subagents and teammates that are not assigned a model another way, set CLAUDE_CODE_SUBAGENT_MODEL=haiku. We have not verified whether a per-model auto-compact window applies inside subagent contexts, so check your usage logs after the first run.
Codex, from the models doc:
codex -m gpt-6-luna
Vercel AI SDK through AI Gateway, from the Vercel changelog:
import { streamText } from 'ai';
const result = streamText({
model: 'anthropic/claude-haiku-5.5',
reasoning: 'low',
prompt: 'Extract the city and date from: Meet me in Paris on May 3.',
});
Routing Table for Worker Jobs#
| Job | Pick | Why |
|---|---|---|
| Exploratory or "go read this" subagent under a Claude orchestrator | Haiku 5.5 | Same price as Luna under 100K, stronger on agentic coding, same wire format as the lead model |
| Bounded code edit or terminal step in a Claude Code or Copilot loop | Haiku 5.5 | 39.2% vs 16.4% Terminal-Bench 4.0 (Anthropic) and 33% vs 13% (Artificial Analysis) |
| Computer-use steps | Haiku 5.5 | 72.4% vs 48.9% on OSWorld 2.1 offline subset, Anthropic-reported |
| Long-running worker that replays 100K-270K tokens of history | GPT-6 Luna | Stays at base price to 272K; Haiku pays 5x past 100K |
| High-volume extraction, classification, structured summaries | GPT-6 Luna, or Haiku at low effort | Luna spends fewer tokens per task; test Haiku at low against your golden set |
| Focused tasks inside Codex | GPT-6 Luna | OpenAI's recommended Codex model for focused, repeatable work |
| Questions that lean on recall of facts | GPT-6 Luna | Higher AA-Omniscience accuracy, though Haiku hallucinates less |
| Anything that needs judgment across many steps | Neither | Use Sonnet 5.5 or GPT-6.1 Sol; see the GPT-6.1 Sol guide |
The routing rule that falls out of all this is simple: route by expected prompt length first and model brand second. A small worker under a frontier orchestrator should get a fresh, short context per task anyway, which is the shape subagents are built for, and in that shape Haiku 5.5 is the better buy at the same price. The moment a worker accumulates history past 100K, Luna is several times cheaper, and no benchmark lead closes a 5x price gap on output.
The Strategic Read#
Anthropic matched Luna's sticker exactly for short prompts, and Anthropic's own footnote says about 90% of Haiku 4.5 requests fell under 100K. That is a targeted entry into the cheap-worker slot where Luna has been the price reference since GPT-5.6 Luna's July price cut, while the 5x long-prompt rate protects Anthropic's margin on the long-context agent traffic where serving cost climbs. OpenAI's structural advantages are runway (272K before any surcharge) and token efficiency; Anthropic's are capability per call and the haiku alias that Claude Code users already route subagents to.
The second-order effect lands on harnesses and gateways. Once two models tie on price but break at different prompt lengths, the router has to know the prompt length before it picks, and the harness has to compact per model rather than per context window. Claude Code already exposes a per-model auto-compact window; expect routers to start treating "how long is this prompt" as a routing input beside "how hard is this task".
What People Are Actually Saying#
The Hacker News launch thread passed 1,000 points and about 480 comments by October 8, and Luna came up constantly.
- The cost-per-task skeptics. One commenter cites Artificial Analysis putting Haiku's cost per task at roughly 3x Luna's at every effort level, and says Luna wins on economics while Sol wins on intelligence. The same commenter adds the case for Haiku as a worker: by their count it is faster, at about 93 tokens per second on OpenRouter so far and at least 137 in Artificial Analysis's runs. Another points out that Artificial Analysis's provisional Haiku costs exclude the over-100K step, which would make the gap wider, not narrower, for long tasks.
- The tokenizer argument. A commenter estimates Claude's 100K tokens equal about 60-65K GPT tokens, so Luna's cutoff is much further away than the headline numbers suggest. That is the reason our worked example carries a +60% column.
- The "stay under 100K" camp. A Luna user says many of their sessions cap out well below 100K and that they will switch if Haiku is that much better.
- The Luna loyalists. Another commenter notes that even at its long-prompt rates Luna is half Haiku's upper tier, and says they are sticking with Luna unless they need a smarter model.
- Cost per completed task. A reply to the Luna loyalist asks whether they are judging by token cost rather than cost per completed task, and says the benchmark in the launch article showed Haiku lower per completed task. That is the right unit for a worker, and it cuts both ways: the Reddit pagoda run above is a single sample of one creative task, not an agent benchmark.
Our read: the community agrees on the facts (same sticker, Haiku smarter, Luna leaner) and splits on workload shape. People running short subagents lean Haiku; people running long, history-heavy loops lean Luna. That is the same line the arithmetic draws.
What We Have Not Tested#
We did not run either model. Every benchmark here is vendor-reported (Anthropic) or third-party (Artificial Analysis), every cost is our arithmetic from published price sheets, and the Reddit numbers are the poster's. We have not measured how Haiku 5.5 and Luna tokenize the same text, which is the single biggest unknown in the cost comparison. If you run both on your own golden set, count tokens with each vendor's tooling and compare cost per completed task, not per million tokens; our cost-per-task explainer has the method.
FAQ#
How do I keep a Haiku 5.5 subagent under the 100K price line?#
Give each subagent call a fresh, short context per task, and in Claude Code set the auto-compact window for Haiku 5.5 to 100K with /autocompact 100k, which the docs save per model. By default Haiku 5.5 sessions compact at about 967K tokens, far past the line where its price rises 5x.
Is Haiku 5.5 a better coding subagent than GPT-6 Luna?#
On published numbers, yes. Anthropic reports 39.2% vs 16.4% on Terminal-Bench 4.0, and Artificial Analysis measured 33% vs 13%. Both are well below Sonnet 5.5, so use either one for bounded steps and subagents, not for whole coding loops.
Why did Haiku 5.5 cost 12x more than Luna in the Reddit test?#
The poster's Haiku run spent most of its tokens on requests over 100K, where Haiku bills 5x. At list prices, the same tokens under 100K would have cost about $5.47, roughly 2.9x Luna's $1.91, and the cliff multiplied that by about 4.6x. The run also used xhigh effort rather than Haiku's medium default.
Which model should I use as a subagent under Claude Opus 5.5 or Sonnet 5.5?#
Haiku 5.5, as long as each subagent call stays under 100K tokens. It matches Luna on price at that size, scores higher on agentic coding, and shares the lead model's API. Set a 100K auto-compact window in Claude Code, or start fresh contexts per task, to keep it under the line.
Does GPT-6 Luna work in Codex and Haiku 5.5 in GitHub Copilot?#
Yes. OpenAI recommends Luna in Codex for focused, repeatable tasks and names it the Free and Go replacement for GPT-5.5. GitHub made Haiku 5.5 generally available in Copilot on October 7 for Pro, Pro+, Max, Business and Enterprise, billed at provider list pricing.
Continue Reading#
- Budget AI Coding Models, October 2026 - the full price comparison of Haiku 5.5, GPT-6 Luna, DeepSeek V4.1-Flash and Gemini 3.5 Flash
- Claude Haiku 5.5 Release Guide - the 100K cliff, the tokenizer change and the ten Haiku 4.5 migration steps
- Opus 5.5 vs Sonnet 5.5 vs Haiku 5.5 - where Haiku sits inside the Claude family and how to route between them
- GPT-6.1 Sol Release Guide - the OpenAI orchestrator Luna usually works under
- Codex Usage Limits and Pricing - Luna's message ranges per five-hour window and its credit rates
- Model Routing Strategies for Cost-Effective Coding - turning a pick-by-task table into a router
- Prompt Caching for the Claude API - the cache math that decides whether Haiku stays cheap
Sources#
| Source | What it supports |
|---|---|
| Anthropic, "Introducing Claude Haiku 5.5" (October 7, 2026) | Benchmark table with the GPT-6 Luna column, positioning, price table, 90%-under-100K footnote, Sonnet and Opus guidance for complex coding |
| Claude Haiku 5.5 model overview | Model ids, 1M context, 128K output, tiered prices including cache rows, Batch discount, default effort, tokenizer note |
| OpenAI, GPT-6 Luna model page | Model id, prices, cache read and write rates, 272K full-request surcharge, Batch and Flex, context, effort range, Chat Completions function-calling limit |
| OpenAI, Codex models | Luna for focused, repeatable tasks, Free and Go replacement for GPT-5.5, gpt-5.4-mini replacement, high starting effort, no Ultra, codex -m gpt-6-luna |
| GitHub Changelog, "Claude Haiku 5.5 in GitHub Copilot" | Copilot availability, plans, surfaces, list-price billing |
| Vercel Changelog, "Claude Haiku 5.5 now available on AI Gateway" | Gateway model name, effort levels, no-markup pricing, AI SDK snippet |
| Claude Code model configuration | haiku alias resolving to Haiku 5.5 (v2.1.293 or later required), claude --model claude-haiku-5-5, 967K default auto-compact, /autocompact range, CLAUDE_CODE_SUBAGENT_MODEL |
| Artificial Analysis, "Claude Haiku 5.5" | Independent Intelligence Index, Terminal-Bench 4.0, AA-Omniscience and hallucination rate, AutomationBench-AA and its over-refusal caveat, output tokens per task, tiered-pricing caveat |
| r/ClaudeAI, "Claude Haiku 5.5 cost 12x more than GPT-6 Luna for the same voxel pagoda" | The 12x claim and the Haiku run's token and cost figures tested above |
| atomic.chat on X (October 7, 2026) | The same comparison, with $24.96 for Haiku 5.5 and $1.91 for GPT-6 Luna |
| Hacker News launch thread | Community reaction, cost-per-task, speed and tokenizer debate |
Get the next comparison like this in your inbox
One email a week on AI Models and the rest of the AI dev stack. Free.
Read next on Claude Code
Budget AI Coding Models, October 2026: Haiku 5.5 vs GPT-6 Luna
Claude Haiku 5.5 and GPT-6 Luna both list at $0.10/$0.50 per MTok. The price tie breaks at Haiku's 100K cliff, and Anthropic's benchmarks favor Haiku. Checked Oct 8, 2026.
11 min readClaude Haiku 5.5: Pricing, Migration Changes, and When to Use It
Claude Haiku 5.5 (claude-haiku-5-5) costs $0.10 input and $0.50 output per million tokens under 100K, with a 1M window and adaptive thinking. The pricing cliff, the tokenizer catch, and the ten migration steps from Haiku 4.5.
9 min readOpus 5.5 vs Sonnet 5.5 vs Haiku 5.5: Which Claude Model to Use
Sonnet 5.5 is the default, Opus 5.5 is for open-ended long-horizon work, and Haiku 5.5 is for bounded steps and subagents. Prices, benchmarks, cost math.
8 min readTechnical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.









