The Text Is a Control Surface

TL;DR
An MCP server tells an agent to run a terminal command it cannot run, and the cost of that sentence grows from 18 points on GPT-5.5 to 69 on GPT-6 Astra. A pipeline loses up to 40.5 points at its own interfaces, and one instruction clause swings 30. Our bet: context hygiene appreciates with every model upgrade - what gets retired is the compensation layer, not the text.
Here is a sentence an agent can read but cannot obey: "Run gcloud auth login to refresh your credentials." The agent has no shell. It has a tool list. The sentence describes an action outside its action space, and the agent, being a good reader, tries anyway.
In one controlled study, an expired-credentials error containing that kind of human-only step left 45 percent of tasks recovered. The same study ran the same errors past five OpenAI models and measured what the unhelpful sentence cost each of them: 18 points on GPT-5.5, and 69 points on GPT-6 Astra. The more capable the model, the more expensive the bad text. That is backwards from the intuition most of us are running on, which is that a smarter model will need less hand-holding, so the strings we feed it matter less over time. The measurements say the opposite. The strings matter more.
In late August we argued that your agent has a five-constraint budget: all-k satisfaction collapses multiplicatively past five or six simultaneous constraints, durable rules die in compaction and handoff, and the fix is a side channel, not a longer prompt. In May we wrote about constraint decay, and in July we argued deep research agents need constraint ledgers. Those pieces were about how many constraints a model can hold and where the durable ones should live. This piece is about the text itself. The strings that reach the model are not documentation and they are not neutral context. They are a control surface, and three fresh measurements from three different threads say the surface's price is going up, not down, as models improve.
The error message was written for you#
The cleanest measurement comes from an audit of MCP error messages across 150 widely used servers. Of 3,001 error messages, 949 tell the caller what to do next, and about half of those next steps depend on something the server cannot see about the caller. On credential errors, 62 of 67 steps ask for a terminal command, a configuration change, or a web page. On rate limits, 20 of 30 say to wait and retry without naming the call to repeat. These messages were written for a human at a keyboard. The caller is an agent with a tool list.
The authors restricted five OpenAI models to the tools available in Berkeley Function Calling Leaderboard tasks and let them hit those errors. The agents did what the step said. On expired credentials, the terminal-command step left 45 percent of tasks recovered, and the loss it caused grew from 18 points for GPT-5.5 to 69 for GPT-6 Astra. GitHub's "Wait before retrying." left 6 percent. The paper's title states the finding plainly: messages written for developers hurt the most capable agents most.
The remedies are cheap and both sides of the interface own one. For MCP server authors, naming a server tool in the step raised recovery on expired credentials to 84 percent (the login tool in place of the shell command) and on rate limits to 88 percent (naming the call to repeat). For agent developers, deleting the human-only step with a one-sentence prompt before the model reads it raised recovery to 82 percent.
So who wins and who loses? Harness and gateway vendors, because a text-normalization layer between servers and models is now a measured product, and anyone who owns a tool surface gets a cheap reliability upgrade by rewriting strings instead of prompts. Who loses? The "just upgrade the model" migration story, which quietly carries a bigger interface bill with every generation. The second-order effect is the interesting one: if interface text is a control surface, then error strings and tool descriptions are load-bearing production artifacts with owners and review, the same way API responses are. The MCP community already half-knows this. On the Hacker News thread about Pi adding MCP support, the most-repeated defense of the protocol is that it is "API plus docs in one place," and an enterprise practitioner notes MCP "has pretty much taken over" integrations. If the docs are the product, the error text is part of the docs, and it is currently written for the wrong reader.
The interface is where the accuracy dies#
The second measurement moves the same argument from error strings to stage boundaries. The decomposition tax holds model, problem, stages, stage prompts, and completion budget fixed, and varies exactly one thing: whether a stage can still see the original problem. The accuracy difference is the tax. Across 21 open-weight models from nine organisations on GSM-Hard and MATH-500, 70 of 118 primary-family tests survive Benjamini-Hochberg correction and 54 survive Holm. The largest cell is 40.5 points, on gemma-3-12B. A placebo that carries at least 60 percent of the extra tokens but at most one word of the problem recovers nothing, so the loss is information, not context length.
Two fixes, both no-model-change, both about where text sits. First, rewriting one stage's instruction moves gemma-3-12B's tax from 4.5 to 36.5 points. Instruction content is not a rounding error; it is a 30-point lever inside a pipeline that already works. Second, adding "every relationship stated between them" to a stage that lists numerical quantities lowered the tax on 9 of 9 models on MATH-500. And the placement finding is the one to tape to your monitor: re-grounding, which shows a stage the original problem again, belongs after the lossy interface. With one lossy interface, re-grounding the stage after it beats re-grounding the stage before it on 7 of 7 models on both benchmarks, and on MATH-500 the earlier repair is worse than none on 7 of 7. Newer models still pay the tax, and the repair still works on them, but the authors' stronger pre-registered rule that predicted which stage would pay was refuted by a sealed held-out test. The practical translation: measure your pipeline one stage at a time, because nobody can tell you where the context dies by looking at the architecture.
This is the kill your agent runs early argument one level down. That piece was about the run lifecycle: kill doomed runs, carry explicit state across the restart. This is the same move at the stage boundary: do not ask a stage to hold what a summary dropped. Re-ground the stage after the loss, keep the relationships when a stage must list numbers, and re-measure whenever you add a stage.
What capability actually retires#
The third measurement is the one that keeps this honest, because it contains the strongest counter-case to everything above. A systematic harness study built a modular NanoHarness on top of mini-SWE-agent and isolated five components across ten open-weight models and three benchmarks. The headline: complex harnesses provide diminishing marginal gains on SWE-style issue repair as model capability improves. On the open-ended repository tasks, that is the sentence every "the next model will fix it" argument quotes.
But read the component-level table, because it splits the budget in two. Structured tool use and task-specific subagents provide the most stable improvements. Context compression and general subagents can hurt repository-generation performance. Combined, the modular harness beats its base by 7.37 percentage points on one model and 6.21 on another, recovering most of the gains of product-level harnesses. So capability does retire something, and it is the compensation layer: the prescriptive tool prompts and general-purpose helpers you bought because the model was weak. What it does not retire is the interface layer: the structure of the tools the model calls and the text it reads to decide.
That is the synthesis, and we have not seen anyone state it this way: the text budget moves in opposite directions under capability growth. Compensation text gets retired by model upgrades. Interface text appreciates. The same upgrade that makes your 400-line system prompt unnecessary makes an unexecutable error message more expensive than it was on the previous generation.
The bet#
Here is the claim, stated so it can be graded. The accuracy cost of the text a model reads is capability-complementary: on a fixed pipeline, the delta from fixing interface text (actionable errors, re-grounding after lossy interfaces, relationship-preserving handoffs) grows with the next model generation, while the delta from compensation scaffolding (prescriptive tool prompts, general subagents) shrinks. Context hygiene is not a thing you mature out of. It is an asset that appreciates at every upgrade.
What would prove us wrong. First, a model generation where unexecutable next-step text stops mattering because the model reads it as data, ignores it, and recovers through its own tool loop; then the MCP loss curve would fall instead of rise, and the whole argument is a snapshot of a transitional era. Second, effect sizes for interface-text fixes that flatline across generations on a fixed pipeline. Third, the protocol layer eating the problem: if MCP or the harnesses auto-translate human-facing error text into callable actions at the transport boundary, per-server text audits stop being necessary and the fix moves to one place. Watch MCP spec changes and SDK lint rules for exactly that. Any of the three would be a real miss, and we will write it up if it lands.
The counter-case, with the steel it deserves#
Start with our own evidence, because it cuts hardest. A 288-run ablation of AGENTS.md files found context-injection strategy did not measurably change coding-agent correctness, bounded to under 10-15 percentage points. If that is the shape of context work in general, why would error strings be different?
Because the claim is narrower than "write more text." The decomposition-tax result says the loss is information, not length, and the MCP result says the cost comes from text that names an action the caller cannot take. Optimized prose that duplicates available source is decoration; a string that names a tool the caller does not have is an anti-instruction. The distinction we now have measurements for is actionable versus decorative, and only the actionable side carries the capability-complementary price. The ablation's null lives on the decorative side; nobody has run the 288-run equivalent with deliberately unexecutable error text, but the MCP study is the closest instrument and it points the other way.
Second, scope. The MCP study is two authors, one benchmark family, and a coded survey that does not execute all 949 messages. The decomposition tax is math-only, single-lab, and "up to 40.5" is a largest cell. The harness study's compression and subagent negatives rest mainly on one benchmark, and its dimensions are not fully crossed. We are reading direction and mechanism, not effect sizes, and the cross-generation compounding in our bet is a collision we assembled, not a single paper's finding. That is why we are labeling the compounding part of this bet speculative.
Third, the practitioner instinct deserves a fair hearing, because it is the strongest form of the counter-case. On the Hacker News thread about keeping Claude Code from forgetting everything, one commenter refuses to invest in configuration at all: none of it will matter when the next model drops, and the special settings of today will be obsolete or backwards tomorrow. On the same thread, others report the file being "mostly ignored" and post-compaction sessions feeling dumber. Both halves are real. The config that compensates for model weakness does expire. The text that defines the tools and the interfaces does not, because it is not compensating for anything; it is the contract. The retirement happens on one side of the line, and the mistake is treating the line as the whole file.
What developers should do#
-
Audit every string the model reads, once per model upgrade. Errors, tool descriptions, skill files, stage prompts. Classify each as actionable (it names something the caller can actually do) or decorative (it duplicates available source or addresses a human). Rewrite the first class, delete the second before it reaches the model. This is a one-hour job that the MCP numbers price at double-digit recovery on auth and rate-limit failures.
-
Name the call, not the human step. If you own a tool surface, replace shell commands, config instructions, and "wait and retry" with the tool name or the call to repeat. If you do not own the server, strip human-only steps in a wrapper before the model reads them. Both remedies measured 82 to 88 percent recovery, versus 6 to 45 percent for the human-facing text.
-
Re-ground after the lossy stage, never before. When a stage summarizes, translates, or compresses, the next stage gets the original problem again. The stage before is the wrong place, and on one benchmark it was worse than no repair at all.
-
Preserve relationships, not just quantities. A stage that lists numbers should be told to keep how they relate. That clause lowered the tax on 9 of 9 models and costs one line.
-
Re-measure harness components at every upgrade, and retire only the compensation layer. Structured tools and task-specific subagents are the stable buys; context compression and general subagents can go negative; complex harness gains shrink as models improve. Treat component value as task-shaped and re-test rather than inherit.
-
Price the text. The same week produced a per-segment cost forecast that beats a fixed token budget by 21.3 percent at matched completion. If you are metering runs at that resolution, meter text at the same resolution: which strings earned their tokens, and which ones were instructions the caller cannot follow.
What people are actually saying#
The threads are split in the way the evidence predicts, with the disagreement landing exactly on the retirement line.
- On the GPT-6.1 Sol thread, practitioners report both directions of literal instruction following. One says the newer model lacks the common sense of the older one and needs more literal prompts; another argues Anthropic models "have always struggled with instructions following." Nobody in the thread thinks text stopped mattering; they disagree about which model reads it best.
- On the Pi.dev MCP thread, the dominant framing is that MCP is "API plus docs," which is the interface-is-the-product position stated from the adoption side. The sharpest dissent is that the protocol is an inferior re-implementation of OpenAPI and that CLI tools with docs work better. Both camps agree the docs are load-bearing; they disagree about the container.
- On the memory and compaction thread, the counter-case gets its best form: do not invest in configuration because the next model obsoletes it. The thread also carries the practitioner report that compaction makes a session "feel so much dumber" and that CLAUDE.md is "mostly ignored" by some agents, which is the interface side refusing to retire quietly.
Continue Reading#
- Your Agent Has a Five-Constraint Budget - the predecessor argument: all-k constraint satisfaction collapses multiplicatively, constraints die in compaction, and durable rules belong in side channels
- Kill Your Agent Runs Early - the state-placement line this piece extends to stage boundaries: kill the run, carry explicit state, re-ground the work
- AGENTS.md Files Don't Move Coding Agent Correctness: A 288-Run Ablation - the measured null that bounds our own claim: context files are behavior steering, and decorative prose is not the lever
- The Response Looked Right. The Work Was Not Done. - why "the model read it" and "the model did it" are different claims, and why the artifact has to carry the check
- The Complete Guide to MCP Servers - the protocol surface where the error-text audit starts, and the place to apply it first
Sources#
- MCP Error Messages Written for Developers Hurt the Most Capable Agents Most - Xiaonan Xu and Wenjing Wu, arXiv:2609.35381, submitted September 28, 2026 (v2 September 29, 2026)
- The Decomposition Tax: LLM Pipelines Lose Up to 40 Accuracy Points at Their Own Interfaces - Tianqi Bu, YuXuan Peng, Junteng Tu, Henghui Xiao, arXiv:2609.32825, submitted September 26, 2026
- Beyond the Model: Demystifying Harness Effects in Software Engineering Agents - Haichuan Hu et al., arXiv:2609.32459, submitted September 26, 2026
- TokenCast: Forecasting Token Consumption During LLM Agent Execution - Chaoqian Ouyang et al., arXiv:2609.35760, submitted September 28, 2026 (v2 September 29, 2026)
- Hacker News: Pi.dev: You Said No MCP - 303 points, 2026-09-30
- Hacker News: GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price - 1001 points, 2026-09-29
- Hacker News: Stop Claude Code from forgetting everything - 202 points, 2025-12-29
Get the next deep dive like this in your inbox
One email a week on Research and the rest of the AI dev stack. Free.
Read next on AI coding tools
Your Agent Has a Five-Constraint Budget
Give an agent one instruction and it obeys. Give it eight and it obeys all of them about five percent of the time, no matter which frontier model you bought. The phase transition is measured, the constraints also die in compaction and handoff notes, and in security-critical code the failure ships as infrastructure. The fix is not a better prompt. It is a smaller simultaneous budget and a side channel for the rules that must survive.
10 min readKill Your Agent Runs Early
The first production-scale trace of agentic coding says the context you keep paying for is already dead at every turn boundary. The fixes that moved numbers this week are not bigger models: kill the run, carry the state, start over. Here is the bet you can grade us on.
11 min readAGENTS.md Files Don't Move Coding Agent Correctness: A 288-Run Ablation
A controlled ablation across Claude Code and Codex, 17 real tasks, and 288 evaluated runs finds context-injection strategy does not measurably change correctness (bounded to under 10-15pp). The failures are implementation skill, not missing repository knowledge.
7 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.





