Claude Advisor Tool vs a Bigger Model: When the Pairing Pays

TL;DR
Anthropic's own runs: a Fable 5.1 advisor added 1.7 points to Opus 5.5 for 2.1x the cost. When the Claude advisor tool pays, and when to raise effort instead.
Use the Claude advisor tool when a cheaper executor runs a long agent loop and actually asks for help; otherwise raise effort or switch models. Anthropic's own coding runs show a Fable 5.1 advisor adding 1.7 points to Opus 5.5 at high for about 2.1 times the cost - roughly what more effort buys.
That is the short version of a cost guide Anthropic has updated with runs from as recently as October 7, and it is less flattering to the advisor pattern than the launch pitch. The advisor tool (advisor_20260301, beta header advisor-tool-2026-03-01) lets an executor model such as Claude Haiku 5.5 or Sonnet 5.5 consult a stronger model mid-request, server-side, in one /v1/messages call. Claude Code exposes the same thing as /advisor, and v2.1.293, the release that added Haiku 5.5, is the minimum version for using Haiku 5.5 on either side of the pairing. The question developers are actually asking is whether to pay for a pairing or just run the bigger model. Here is what the primary sources say, with every number labelled.
Official Sources#
| Resource | What it confirms |
|---|---|
| Advisor tool docs (Claude Platform) | Tool type, beta header, parameters, compatibility table, billing, prompting guidance |
| Optimizing for cost and intelligence | Anthropic's measured advisor pairings, consult rates and cost per attempt |
| Claude Code advisor docs | /advisor, --advisor, advisorModel, accepted pairings and version requirements in Claude Code |
| Claude API pricing | Per-MTok rates for every executor and advisor model below |
| Claude Code v2.1.293 release | Haiku 5.5 added, $0.10/$0.50 per MTok ($0.50/$2.50 over 100K) |
| Claude Code v2.1.290 release | Advisor calls surfaced to mods as serverToolUses |
At a Glance: Executors, Advisors, Prices#
The rule from the API docs: the advisor must be Claude Sonnet 4.6 or more capable, and at least as capable as the executor. An invalid pair returns a 400 invalid_request_error. Prices are list input/output per million tokens from Anthropic's pricing page, checked October 9, 2026.
| Executor (price in / out) | Current-generation advisors it accepts | Advisor prices in / out |
|---|---|---|
Haiku 5.5 (claude-haiku-5-5), $0.10 / $0.50 up to 100K prompt | Haiku 5.5, Sonnet 5, Sonnet 5.5, Opus 4.7, Opus 4.8, Opus 5, Opus 5.5, Fable 5, Fable 5.1, Mythos 5, Mythos 5.1 | Sonnet 5.5 $2 / $10, Opus 5.5 $4 / $20, Opus 5 $5 / $25, Fable 5.1 $10 / $50 |
Sonnet 5.5 (claude-sonnet-5-5), $2 / $10 | Sonnet 5.5, Opus 5, Opus 5.5, Fable 5, Fable 5.1, Mythos 5, Mythos 5.1 | Opus 5.5 $4 / $20, Fable 5.1 $10 / $50 |
Opus 5.5 (claude-opus-5-5), $4 / $20 | Opus 5, Opus 5.5, Fable 5, Fable 5.1, Mythos 5, Mythos 5.1 | Fable 5.1 $10 / $50, Mythos 5.1 $10 / $50 |
Fable 5.1 (claude-fable-5-1), $10 / $50 | Fable 5.1, Mythos 5.1 | $10 / $50 |
Two things in that table surprise people. Sonnet 5.5 no longer accepts a Sonnet 5 or any Opus 4.x advisor, which is one of the breaking changes in the Sonnet 5.5 release guide. And the API docs list the advisor as available in beta on the Claude API and Claude Platform on AWS, but not on Amazon Bedrock, Google Cloud or Microsoft Foundry. Claude Code's own advisor page is stricter and says it needs the Anthropic API, listing Claude Platform on AWS among the unsupported routes. If you run on a cloud marketplace, the decision is already made for you: raise effort or switch models.
What One Advisor Call Costs#
Anthropic's docs say advisor output is "typically 400 to 700 text tokens, or 1,400 to 1,800 tokens total including thinking", and the advisor reads the executor's full transcript as input. The worked numbers below are our arithmetic at list prices, assuming a 50,000-token transcript, 1,600 advisor output tokens, and advisor-side caching off (its default).
| Advisor | Input (50K tokens) | Output (1,600 tokens) | One call |
|---|---|---|---|
| Haiku 5.5 | $0.005 | $0.0008 | about $0.006 |
| Sonnet 5.5 | $0.10 | $0.016 | about $0.12 |
| Opus 5.5 | $0.20 | $0.032 | about $0.23 |
| Opus 5 | $0.25 | $0.04 | about $0.29 |
| Fable 5.1 | $0.50 | $0.08 | about $0.58 |
The pattern that falls out: every model in this table prices output at five times input, so once the transcript is more than five times the length of the advice, the advisor's bill is mostly its reading, not its writing. On a long coding loop that is almost always true. Two levers follow, both from the docs:
caching: {"type": "ephemeral", "ttl": "5m"}on the tool definition caches the advisor's own transcript across calls. Anthropic says it "breaks even at roughly three advisor calls", so turn it on for long loops and leave it off for short tasks. Claude Code's advisor page says Claude Code does not cache the advisor's read at all: each call processes the full transcript anew.max_tokenson the tool definition (minimum 1024) caps the advisor's thinking plus text per call. On a hard reasoning benchmark, Anthropic measuredmax_tokens: 2048cutting mean advisor output from about 4,200 to 5,900 tokens down to about 630 to 840, with near-zero truncation. The top-levelmax_tokensdoes not bound advisor tokens.
There is also a cost you pay even when the advisor is never called. The tool definition is "about 1,000 prompt tokens per request". On GPQA Diamond, where neither executor consulted once, that alone added 12% to Haiku 5.5's cost per question and 25% to Sonnet 5.5's. For the economics of caching that prefix, see our Claude prompt caching production guide.
Benchmarks: What Anthropic Measured#
All rows are vendor-reported by Anthropic on its Optimizing for cost and intelligence page. The coding benchmark is Anthropic-internal (370 repository tasks graded by the repositories' own tests), so it cannot be reproduced. No third party has published advisor-pairing measurements we could find.
| Configuration | Workload | Consult rate | Result | Cost |
|---|---|---|---|---|
Opus 5.5 at high + Fable 5.1 advisor | Internal agentic coding | 1.39 requested, 1.35 received per attempt | 90.1% over five attempts per task, +1.7 points over Opus 5.5 alone at high ("at the edge of run-to-run noise"); the 279 attempts in which the advisor was turned away under load were re-run | $2.92 per attempt vs $1.38 alone |
Opus 5.5 alone at xhigh | Internal agentic coding | n/a | 91.1% (one attempt per task) | $4.11 per attempt |
Opus 5.5 alone at medium (default) | Internal agentic coding | n/a | 86.6% | $0.84 per attempt |
Fable 5.1 alone at medium | Internal agentic coding | n/a | 84.2% (single run) | $2.68 per attempt |
Fable 5.1 alone at high (its default) | Internal agentic coding | n/a | 85.7%, 317 of 370 tasks (one attempt per task) | Not stated |
| Haiku 5.5 + Opus 5.5 advisor | GPQA Diamond | None of the 198 questions in either of two runs (0 of 396) | Within noise of Haiku 5.5 alone (85%) | +12% from the tool definition |
| Sonnet 5.5 + Opus 5.5 advisor | GPQA Diamond | None of the 198 questions in either of two runs (0 of 396) | Within noise of Sonnet 5.5 alone (91%) | +25% from the tool definition |
Opus 5.5 at low + Fable 5.1 advisor | Chartography (chart reading) | 1 of 300 tasks | 61.7, 7 points below Opus 5.5 alone (68.7) | About the same cost |
| Sonnet 5 at low effort + advisor | DeepSWE | Kept asking | +23 points | Client-side advisor loop, not the tool |
One qualifier on the headline coding result: Anthropic's footnote says the 279 pairing attempts in which the advisor was turned away under load were re-run, while attempts whose consults timed out were kept. The comparison point Anthropic charts it against, Fable 5.1 alone at high, solved 317 of 370 tasks (85.7%) with one attempt per task.
Read the table as two findings. First, the advisor buys capability only when the executor asks. Three of the four pairings on Anthropic's chart rarely or never consulted, and Anthropic's GPQA runs added the tool "without a system prompt that asks the executor to consult it". Second, when it does work, the advisor lands on the executor's own effort curve: Anthropic says it "buys about what more effort does". Opus 5.5 alone at xhigh scored a point higher than the pairing (from one attempt per task, against five for the pairing), for $4.11 instead of $2.92 per attempt, but with no beta header and no second model to manage.
The cost mechanism is visible in the coding numbers. With an Opus 5.5 executor, the advice saved almost nothing on the executor side ($1.36 per attempt against $1.38 alone) while the consultations cost $1.55. In August, with an Opus 5 executor, advice saved $1.26 per attempt in executor tokens against $2.47 of consultations, paying back about half. That is the "an executor told the right approach explores fewer dead ends" effect, and it shrinks as the executor gets closer to the advisor.
Pick by Task#
| Situation | Pick | Why |
|---|---|---|
| Haiku 5.5 doing classification, extraction, routing or summaries | No advisor | Nothing to plan, and the tool definition alone added 12% per call on Haiku 5.5 in Anthropic's GPQA runs |
| Haiku 5.5 or Sonnet 5.5 running a long multi-step coding agent | Opus 5.5 advisor plus the docs' coding system prompt | Most turns stay at executor rates; the prompt is what gets the executor to consult |
| Opus 5.5 coding agent that needs a few more points | Raise effort to high or xhigh first | The Fable 5.1 pairing landed on Opus 5.5's own effort curve, with no beta dependency |
Any executor at low effort | No advisor | Anthropic saw a low-effort executor stop consulting and score 7 points below the executor alone |
| Single-turn Q&A or a user-facing model picker | No advisor | The docs call these a weak fit: nothing to plan, or users already choose the trade-off |
| Running on Bedrock, Google Cloud or Foundry | Effort or model switch | The advisor tool is not available there |
| Claude Code with Sonnet as the main model | /advisor opus | Anthropic's suggested pairing: Sonnet does routine work and escalates planning, ambiguous failures and completion checks |
| High-stakes change where an independent check matters more than cost | Opus main + Opus advisor in Claude Code | The Claude Code docs list it for that case; expect to pay for both |
The baseline rule from Anthropic is worth taping to the monitor: "first price the advisor's model alone at low effort; that is the baseline to beat." If the executor consults on most tasks, you are paying advisor rates across the whole workload and running the advisor's model directly is cheaper. For the older two-dial version of this decision, effort versus switching models, see Fable 5 effort vs model switching.
Setting It Up#
Claude API#
This is the docs' quick start with the executor swapped for Haiku 5.5, the advisor for Opus 5.5 (a valid pair in the compatibility table), and the two cost controls added. We did not run it: no API key was available in this environment.
import anthropic
client = anthropic.Anthropic()
response = client.beta.messages.create(
model="claude-haiku-5-5", # executor
max_tokens=4096,
betas=["advisor-tool-2026-03-01"],
tools=[
{
"type": "advisor_20260301",
"name": "advisor",
"model": "claude-opus-5-5", # advisor, billed at its own rates
"max_tokens": 2048,
"caching": {"type": "ephemeral", "ttl": "5m"},
},
# ... your other tools
],
messages=[{"role": "user", "content": "Build a concurrent worker pool in Go with graceful shutdown."}],
)
The gotchas, all from the API docs:
- Round-trip results verbatim. Some advisors return an encrypted
advisor_redacted_result(the docs' example is Opus 5) and some return plaintextadvisor_result(Opus 4.8). Pass the block back unchanged on the next turn, and branch oncontent.typeif you switch advisors. - Top-level
usageexcludes advisor tokens. Sumusage.iterationsand priceadvisor_messageentries at the advisor's rates, or your cost dashboard will under-report. - Streaming goes quiet. The advisor sub-inference does not stream; you get SSE pings roughly every 30 seconds, then the whole result in one event.
pause_turnneeds a resend with the same advisor tool and beta header, or the API returns a 400.- You cannot force a consult on the newest executors. Opus 5.5, Sonnet 5.5, Fable 5.1 and Mythos 5.1 executors reject
tool_choicetypestoolandany, so use a prompt nudge. On Haiku executors, the docs' turn-2 reminder raised pass rates by "roughly 7 percentage points"; on Opus it slightly lowered them. - Budgets do not cover it. Advisor tokens do not draw from a task budget on the executor, and a Priority Tier commitment on the executor does not extend to the advisor.
If your agent already juggles client and server tools, the tool use production patterns guide covers the loop shape this slots into.
Claude Code#
Three ways to turn it on, verbatim from the Claude Code docs:
/advisor opus # in a session; saved to advisorModel
claude --advisor opus # one session only
/advisor off # turn it off
Or set "advisorModel": "opus" in your settings file, and CLAUDE_CODE_DISABLE_ADVISOR_TOOL=1 disables it entirely. Haiku 5.5 as main model or advisor needs Claude Code v2.1.293 or later; Fable 5.1 needs v2.1.257 or later. There is no setting to cap or force calls, so if you want more or fewer consultations, say so in your instructions.
The trap worth knowing: if the API refuses a pairing Claude Code attached, Claude Code resends the request without the advisor, and "you see no error and get no advisor calls". The advisor is also fetched behind a feature flag, so a session with DISABLE_TELEMETRY set keeps it off. If you turned it on and the transcript never shows an Advising line, check both before concluding the model just did not need help. Before adding a second model, our Opus 5.5 Claude Code playbook covers how to steer the main model itself.
What People Are Actually Saying#
- The UX case is real for interactive work. On an Ask HN thread about default models, wdm0006 runs "sonnet as main w/ opus advisor" in Claude Code and finds "a fast main model with a big smart advisor is a nice UX" when at the keyboard, but cares more about cost and quality than speed for unattended agents. That matches Anthropic's numbers: the advisor buys responsiveness and a few points, not a free lunch.
- Power users stack it with model routing. On the Opus 5.5 thread, rdli describes "a setup that uses Fable 5.1 as an advisor, Sonnet 5.5 for mechanical changes", the frontier-advisor-over-mid-tier-executor shape Anthropic calls most cost-effective.
- The counter-case: agreement is not verification. On a thread about LLM judges agreeing, ex1fm3ta finds it "kinda funny" that the Claude Code advisor "agrees with the ideas that the previous model did". The advisor sees only the executor's transcript, so it inherits the executor's framing. It is not a substitute for tests. LivePlan's corrective-steering research takes a different tack, letting a monitor decide when the advisor gets called instead of the executor.
- Open harnesses want the same thing. In the local Qwen discussion, walthamstow notes "Claude kind of has this already in their Advisor feature" and that open harnesses could call out to bigger models the same way. Our write-up of that thread is Local Qwen is a different tool, not a worse Opus.
FAQ#
What is the Claude advisor tool?#
It is a server-side tool (advisor_20260301, beta header advisor-tool-2026-03-01) that lets an executor model consult a stronger advisor model mid-request. The executor decides when to call it, Anthropic runs the advisor on the full transcript, and the advice comes back as an advisor_tool_result block in the same /v1/messages response.
Is the advisor tool cheaper than just using Opus?#
Sometimes. Anthropic's docs say a Sonnet executor with an Opus advisor "keeps total cost similar or lower" than Sonnet alone on complex tasks, but its own Opus 5.5 plus Fable 5.1 coding pairing cost about 2.1 times Opus 5.5 alone at high for a 1.7-point gain. If the executor consults on most tasks, running the stronger model directly is cheaper.
Which advisor should I pair with Haiku 5.5?#
Opus 5.5 is the natural pick at $4/$20 per million tokens, and the API accepts everything from Haiku 5.5 up to Mythos 5.1 as its advisor. Use the docs' coding system prompt or turn-2 nudge, because in Anthropic's GPQA runs a Haiku 5.5 executor with an Opus 5.5 advisor never consulted it.
How do I turn on the advisor in Claude Code?#
Run /advisor opus in a session, launch with claude --advisor opus, or set "advisorModel": "opus" in settings. Haiku 5.5 as the main model or advisor needs Claude Code v2.1.293 or later, and /advisor off turns it off.
Does the advisor tool work on Bedrock or Vertex?#
No. The API docs list it in beta on the Claude API and Claude Platform on AWS only, not on Amazon Bedrock, Google Cloud or Microsoft Foundry. Claude Code's advisor page says it requires the Anthropic API.
Sources#
| Source | URL |
|---|---|
| Anthropic: Advisor tool documentation | platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool |
| Anthropic: Optimizing for cost and intelligence (advisor strategy measurements) | platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence |
| Anthropic: Claude Code advisor documentation | code.claude.com/docs/en/advisor |
| Anthropic: Claude API pricing | platform.claude.com/docs/en/about-claude/pricing |
| GitHub: Claude Code v2.1.293 release notes | github.com/anthropics/claude-code/releases/tag/v2.1.293 |
| GitHub: Claude Code v2.1.290 release notes | github.com/anthropics/claude-code/releases/tag/v2.1.290 |
| Hacker News: Ask HN, what default model do you use and why? | news.ycombinator.com/item?id=49672966 |
| Hacker News: Getting the most out of Opus 5.5 in Claude and Claude Code | news.ycombinator.com/item?id=49946567 |
| Hacker News: When LLM judges agree, should we believe them? | news.ycombinator.com/item?id=49699590 |
| Hacker News: Local Qwen isn't a worse Opus, it's a different tool | news.ycombinator.com/item?id=48580209 |
Last updated: October 9, 2026
Continue Reading#
- Claude Opus 5.5 Developer Guide: Pricing, API, Claude Code - the executor and advisor most pairings above are built around
- Claude Haiku 5.5: Pricing, Migration Changes, and When to Use It - the cheapest executor, and the 100K pricing cliff to watch
- Prompt Caching in the Claude API: A Production Guide - the lever that matters more than any second model
- Tool Use in the Claude API: Production Patterns for Reliable Agents - the agent loop the advisor plugs into
- Fable 5 Effort vs Model Switching - the effort-first argument applied to Fable
Get the next comparison like this in your inbox
One email a week on Claude and the rest of the AI dev stack. Free.
Read next on Claude Code
Claude Opus 5.5 Developer Guide: Pricing, API, Claude Code
Claude Opus 5.5 costs $4/$20 per million tokens with $0.20 cache reads. SDK examples, Claude Code setup, and when to pick it over Sonnet 5.5 or Haiku 5.5.
11 min readClaude Haiku 5.5: Pricing, Migration Changes, and When to Use It
Claude Haiku 5.5 (claude-haiku-5-5) costs $0.10 input and $0.50 output per million tokens under 100K, with a 1M window and adaptive thinking. The pricing cliff, the tokenizer catch, and the ten migration steps from Haiku 4.5.
9 min readPrompt Caching in the Claude API: A Production Guide
Cut Claude API spend by up to 90% with prompt caching. Real numbers, TypeScript SDK code, and the gotchas Anthropic's docs gloss over.
11 min readTechnical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.







