Opus 5.5 in Claude Code: The Playbook and Where It Breaks

TL;DR
claude.dev's Opus 5.5 playbook says to name the finish line, delete 'think hard' lines and write stop rules into CLAUDE.md. What each tip changes, plus what 130 Hacker News comments say about where long runs go wrong.
The claude.dev blog published a playbook for Opus 5.5 in Claude and Claude Code, and it reached the Hacker News front page on October 3. Its core advice: give Opus 5.5 the whole task in one message, say what "done" looks like, delete "think carefully" from your prompts, and write a short rule in CLAUDE.md for when it should keep going and when it should stop and ask. Below: what each tip changes in practice, the one setting that costs money, and where the community says long runs still go wrong.
Last updated: October 3, 2026
What the playbook says to change#
The guide by Addy Osmani was published September 22 and is about nine minutes long. It is vendor guidance, so read it as the intended way to use the model, not as a measurement. The changes that matter for Claude Code:
| Change | What to do | Why the guide says it matters |
|---|---|---|
| Name the finish line | Put the whole task in one message with a testable end state, such as "the tests pass" or "every endpoint is migrated" | Opus 5.5 keeps going on multi-step work better than Opus 5, and a clear end state tells it when to stop |
| Drop "think hard" | Delete "think carefully" and similar lines from prompts and saved instructions | The model thinks before every reply and decides how much. In the authors' chat testing, removing the line made replies start sooner with no clear quality drop |
| Say when to stop | Add a CLAUDE.md rule: keep going when a step needs no input, stop for anything destructive | On long tasks it sometimes stops to report or offers to continue instead of acting |
| Fan out to subagents | For audits and migrations, ask it to give each unit to its own subagent and check each result's evidence | Early testers had it coordinate parallel subagents with little oversight |
| Keep the task list in a file | Ask for a checklist file it updates as it goes | A long run fills the context window and older turns get summarized; a file survives that |
| Review before a person does | Ask it to list only merge-blocking problems with file, line and a way to show the failure | One early tester reported that Opus 5.5 at its lowest effort caught more bugs than Opus 5 at high effort |
| Mark what it could not confirm | Add "mark anything you couldn't confirm, and say where you looked" | Makes gaps easy to find in research and analysis |
For design work the guide recommends naming the styles you do not want rather than asking for "not generic", because a specific list changes the output more than a general instruction.
A CLAUDE.md stop rule to start from#
The guide's own example is two sentences. This is our adaptation of the idea, not a quote, so edit it to fit your repository:
When a step does not need my input, keep going and put status notes in the
same message as your next action. Stop and ask only when you cannot continue
without me, or before anything destructive: deleting data, force-pushing, or
changing anything outside this repository.
Keep your permission prompts on for destructive commands as well. The guide says the same, and the Claude Code permissions and settings guide covers how to scope them. If you run parallel agents, our subagents vs agent teams vs workflows comparison helps pick the shape, and the CLAUDE.md memory failure post is worth reading before you pile rules into one file.
The setting that costs money: fast mode#
The guide lists fast mode for Opus 5.5 as a research preview at launch: same model, text arrives sooner, but it needs extra usage turned on and costs more per token than standard mode. Turn it on with /fast for back-and-forth work where you read each reply. For long unattended runs it buys little. Our fast mode cost breakdown has the per-token math.
Safeguards can switch your model mid-session#
Easy to miss: the guide says Opus 5.5 launched with new bio and cyber safeguards, and in Claude Code most flagged messages move the session to an older model instead of stopping it. The check covers everything in the conversation, including files and search results, so a flag can come from earlier content. To go back, run /model. To be asked before a switch, change the setting in /config. The guide says the safeguards are being tuned to cut incorrect flags, and /feedback is the route if a flag was wrong. Also, the guide says asking the model to reproduce its internal reasoning in the reply can be declined, so ask for a short explanation of the choice instead.
What people are actually saying#
The Hacker News thread had about 130 points when we read it on October 3. The tone is mostly positive, with specific caveats that matter for the tips above:
- Long runs work for some. One commenter handed it general CI-speed directives, asked it to plan and have a second model check the plan, and reported 12 PRs ready to merge nine hours later. Another calls it the first model they trust for tasks over an hour.
- Independence cuts both ways. One commenter found it too interested in acting alone, making calls against their stated recommendation, and says it pushed an auto-mode permission well past what was authorized. That is the case for the guide's "stop before anything destructive" rule, and for scoping permissions tightly rather than relying on the prompt.
- Not every workload suits long tasks. A skeptic says long runs waste time on bad decisions that one clarifying question would have avoided and asks which workloads actually benefit. The guide's answer is a clear finish line and stop conditions; the thread does not settle it.
- Plan limits are a recurring complaint. Several commenters say the $20 plan goes a long way, while one asks for a middle tier between $20 and $100.
These are individual reports, not measurements. None of them includes a benchmark.
Where it breaks#
- Vendor guidance. The guide's claims about testers and quality come from its authors. We have not reproduced them, and nothing here is a test of Opus 5.5.
- "Keep going" rules trade stops for risk. The more you tell it to continue, the more your protection depends on permission settings and the destructive-action stop.
- Removing "think hard" is a chat-product finding. The guide reports it from chat testing. In Claude Code, effort is the setting that changes how much it thinks.
- Flag-driven model switches can change behavior mid-task, so check which model a long run ended on.
FAQ#
Should I delete "think step by step" from my prompts for Opus 5.5?#
The guide says yes: Opus 5.5 thinks before every reply, and removing a "think carefully" line made replies start sooner without a clear quality drop in the authors' chat testing. To change how much it thinks in Claude Code, change the effort setting instead.
What should a CLAUDE.md rule for long runs say?#
Say when to keep going (steps that need no input) and when to stop (cannot continue without you, or anything destructive). Keep permission prompts on for deleting data and force-pushing too.
How do I stop Claude Code switching away from Opus 5.5?#
Run /model to switch back, or change the "switch models when a message is flagged" setting in /config so Claude Code asks first. Run /feedback if the flag was wrong.
Is fast mode worth it with Opus 5.5?#
It suits interactive work where you read each reply before sending the next. It costs more per token and needs extra usage enabled, so it is a poor fit for unattended runs.
Continue Reading#
- Claude Opus 5.5 Developer Guide - API examples, Claude Code setup, pricing and when to use it
- Claude Code Permissions and Settings Guide - how to scope the destructive-action stop
- Is Claude Code Fast Mode Worth It? - the cost side of the one paid setting
- Claude Code Subagents vs Agent Teams vs Workflows - which shape fits a large audit or migration
- Claude Code Tips and Tricks - the broader habit list
Sources#
| Source | URL |
|---|---|
| claude.dev: Getting the most out of Opus 5.5 in Claude and Claude Code (Addy Osmani, September 22, 2026), fetched October 3, 2026 | https://claude.dev/blog/getting-the-most-out-of-opus-5-5/ |
| Hacker News discussion, read October 3, 2026 | https://news.ycombinator.com/item?id=49946567 |
| Anthropic: Claude Opus 5.5 | https://www.anthropic.com/claude-opus-5-5 |
Get the next deep dive like this in your inbox
One email a week on Claude Code and the rest of the AI dev stack. Free.
Read next on Claude Code
Claude Opus 5.5 Developer Guide: API Examples, Claude Code Setup, Pricing, and When to Use It
Claude Opus 5.5 (claude-opus-5-5) is Anthropic's new default Opus: $4/$20 per million tokens, $0.20 cache reads, 1M context, thinking always on with medium default effort. Runnable TypeScript and Python SDK examples, Claude Code setup, before/after prompts, and a decision guide vs Sonnet 5, Haiku 4.5, and Fable 5.1.
11 min readClaude Code Permissions: settings.json Allow, Deny, Ask
Configure Claude Code permissions in settings.json: allow, deny, and ask rules, scope precedence, tool specifiers, and the headless flags you need for CI.
11 min readClaude Code Fast Mode: When 2.5x Speed Is Worth 2x Price
Claude Code fast mode pricing explained: $10/$50 per MTok on Opus 4.8, the first-enable context charge, separate rate limit pools, and when 2.5x speed pays off.
8 min readTechnical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








