
TL;DR
Twelve frontier models sat at 60 percent on scientific coding, successors tying predecessors - a textbook saturation curve. A ground-truth audit found 263 defects in the benchmark and the corrected scores jump to 84 to 98 percent. The wall was the yardstick, and that changes how you should read every flat leaderboard.
Twelve frontier model snapshots, all parked between 45 and 60 percent on scientific coding, and the newest ones tying the ones before them. That is the shape of a wall. It is also the shape of a broken yardstick, and the audit that separates the two is the strongest single number we have seen in months of reporting on benchmark integrity.
SciCode-Verified did what nobody else had done with the most-cited scientific coding benchmark: a domain-expert audit of all 65 problems, line by line (arXiv:2608.04975). They found 263 defects. 192 of them, across 91 percent of the main problems, wrongly reject correct, instruction-following solutions - through non-reproducible gold answers, over-tight tolerances, and self-contradictory specs. After correcting every confirmable defect, twelve frontier model snapshots jump from 45-60 percent to 84-98 percent on subproblem accuracy, and from 9-27 percent to 69-92 percent on main problems. The tight cluster of 2026 models around 60 percent, the "successors tying predecessors" reading that fuels saturation narratives, dissolves. It was never a capability plateau. It was 263 benchmark bugs.
In early August we argued that your benchmark is lying to you - that the gap between what agent benchmarks report and what actually happened is routinely double-digit, and systematic rather than random (your-benchmark-is-lying-to-you). We set the floor at "any delta under about 15 points is indistinguishable from measurement error," and then argued the fixes would be architectural, not model-side: ledgers, counterfactuals, deterministic verdicts (the-benchmark-fix-is-architectural). This audit is that position coming home. But it also takes it somewhere new, and we want to defend the new claim properly: the most dangerous benchmark number is not a wrong delta between two models. It is a flat curve, because a flat curve reads as science and gets funded like one.
A wrong delta costs you a bad model choice. A wrong plateau costs you a strategy. When a benchmark reports that successors tie predecessors, the field reads it as "we have hit the wall" - research budgets get reallocated, scaling narratives get rewritten, and the reading quietly becomes a policy position. It looks different from noise because it has the shape of a result: consistent, monotone, stable across model generations.
SciCode-Verified is the cleanest falsification of that reading we have. The benchmark is not obscure: it is part of the Artificial Analysis Intelligence Index and standing government and lab suites. It was reporting saturation on the exact axis - scientific coding - where the plateau narrative is loudest. And 78 percent of the score-suppressing defects required specialized physics and math knowledge to detect. Not clerical proofreading. Not an LLM judge. A working physicist reading a tolerance and realizing it cannot be satisfied.
That last number matters more than the score flip. It means the plateau was not removable by the standard eval-hygiene toolkit. Nobody's cheap script was going to find it, which is exactly why it survived so long, and why the correction needed a domain expert to be the instrument.
The same scout batch delivered a second structural result, from the other end of the measurement stack. A fully-crossed study of three instruction-tuned models, five inference frameworks, six benchmarks, and four generation modes found that the serving backend is a non-negligible factor in measured performance even under greedy decoding, where sampling noise is eliminated by design (arXiv:2608.04714). Roughly 39 percent of the variability a practitioner sees out of the box can trace to the serving framework rather than the model, with the rest from sampling noise and per-framework defaults. The divergences are worse on factual benchmarks than on social-bias ones.
The conclusion is a discipline change, not a curiosity: a model does not have a number. A (model, backend, version, generation-configuration) tuple has a number. Every published score that omits any of those four is incomplete, and every "model X beats model Y by 3 points" headline that does not name the serving stacks is unfalsifiable as published. This also quietly re-prices our own fleet-economics content: if a comparison shows the cheap tier at a fraction of the cost, check whether you are comparing models or comparing backend defaults. Cheaper can be a serving configuration, not a model.
From the archive
Aug 6, 2026 • 6 min read
Aug 5, 2026 • 7 min read
Aug 5, 2026 • 7 min read
Aug 5, 2026 • 8 min read
Then there is the one that reads like a trick and is a measured result. MirageBench ran 12 models across 7 families on 150 personas and 6 personalization tasks, judged 143,616 claims with an independent judge validated against blind human annotation at kappa 0.863 (arXiv:2608.04570). Every single model over-inferred: fabricated user attributes beyond the evidence on 35-49 percent of its claims, mean 41.6 percent. Inferred attributes accumulate nearly linearly across turns with little revision - early guesses harden into profile facts.
And the flagship finding is an inversion: at the model-selection level, a model's self-assessment of its own over-inference is negatively rank-correlated with the judge-measured rate (rho = -0.60, p = 0.044, wide CI at n = 12). The models that report the least fabrication are flagged as fabricating the most. The vendor line "our model is less presumptuous" can be exactly backwards, and there is no self-report that rescues you from checking.
We covered the shape of this before: the judge leaves the loop because LLM verdicts cannot be trusted to grade themselves (the-judge-leaves-the-loop). This is the same lesson with a new victim. Profile inference is memory, and memory that infers is memory that fabricates - which is why any personalization surface or agent memory store needs write-path validation of the inference step, not just the storage step. The external store survives as verification and hygiene, never as retrieval cleverness.
None of this means every plateau is a broken yardstick, and we want the honest boundary drawn before someone quotes this piece the wrong way.
First, audits are expensive. 78 percent of the SciCode defects needed domain experts to find, and the audit is one benchmark of 65 problems. Nobody is auditing every benchmark in every release cycle, so for most flat curves you will not get the falsification. The correct reading is not "plateaus are always instruments." It is "a plateau is a hypothesis about the instrument before it is a fact about the models," and the burden of proof sits on the people quoting the wall.
Second, corrected benchmarks can overcorrect. The corrected SciCode numbers cluster at 84-98 percent, which is itself suspiciously tight - the same instruments that produced a false plateau could be compressing the top of the distribution in the other direction. The authors re-checked every correction independently, which is more than most audits do, but one audit is one audit. We grade our own claims, and the honest grade here is: the plateau reading is dead; the corrected ordering is provisional.
Third, some walls are real. Our own oncall thesis rests on ORCA-bench Hard sitting around 10 percent across frontier agents - an instrument that has been getting harder, not easier, under every check we have seen (swe-nfi-coding-agents-quality-benchmark). The point is not that capability never plateaus. The point is that the reading is not self-validating, and the SciCode case proves the cost of treating it as such.
Here is the claim, stated plainly so it can be graded: saturation readings on frontier benchmarks are instrument hypotheses until audited, and the plateau narrative - successors tying predecessors - is the single most consequential form of eval noise because it looks like a result. By end of 2027, published benchmark claims will carry audit metadata as routine (ground-truth validation rates, failure-cause breakdowns), and the specific tell this run exposed - a flat, stable cluster across model generations - will be the trigger that gets a benchmark audited rather than quoted.
What would prove us wrong: a corrected frontier benchmark whose plateau survives the audit, with the corrections themselves independently re-checked. We have exactly one clean case in our favor so far. We need more, and we will report the failures with the same care as the wins, because a grading desk that only publishes the favorable audits is just another benchmark with a bug in its ground truth.
Treat flat as suspicious, not reassuring. When the newest model ties the last one on a headline benchmark, that is the moment to ask whether the instrument has been audited - not the moment to conclude the field stalled. The saturation reading is the most expensive number on the page.
Demand the tuple, not the number. Every score you act on should come with model, backend, version, and generation configuration. If a vendor or a paper gives you one number and no serving details, the number is a claim, not a measurement.
Never buy "less presumptuous." If a model or product self-reports low fabrication or high honesty, treat it as an unverified claim, because the measured cross-model pattern is inversion, not correlation. Ask for the external judge, the kappa, the strata. Vendors who only have the self-report have told you something too.
Audit the inference step, not just the store. If your agent or product keeps inferred user attributes, those attributes are ~40 percent fabricated by default and they harden across turns. Version the profile, validate writes, and expect the inference to lie in the direction of confidence.
Re-run the frontier math on corrected numbers. The measured frontier gap is partly measurement: corrupted ground truth, undisclosed backends, and missing verified-handoff architecture all inflate the spread. A cheap scout with sandbox-verified notes plus a frontier fixer ties the best single model at a fifth of the cost per solve on one SWE-bench Pro slice (arXiv:2608.04804). Before you pay the premium, ask what the spread looks like on an audited, backend-disclosed instrument.
Read next
A wave of audits in the last two days measured the noise floor of agent benchmarks: misaligned ground truth, lenient model judges, and aggregate scalars that hide real failures. Here is what the numbers actually mean, what to trust, and how to buy agents without being played.
10 min readSince we published your-benchmark-is-lying-to-you, roughly 25 new results have landed on the eval-integrity question. The surprise: every fix that works is structural - ledgers, counterfactuals, decompositions, personas, readout discipline - and none of them asks the model to be smarter.
11 min readEvidence gates, verifiable reward games, deploy-time certificates: the fixes that moved agent quality this week did not make judges better, they removed the judge. We think the LLM verdict inside the agent loop is a transitional technology, and here is the bet you can grade us on.
10 min readTechnical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.
Set up Codex Chronicle on macOS, manage permissions, and understand privacy, security, and troubleshooting.
Getting Started2.5x faster Opus at a higher token cost (research preview).
Claude CodeResearcher, auditor, reviewer, and other ready-made subagent types.
Claude Code
A wave of audits in the last two days measured the noise floor of agent benchmarks: misaligned ground truth, lenient mod...

Since we published your-benchmark-is-lying-to-you, roughly 25 new results have landed on the eval-integrity question. Th...

Evidence gates, verifiable reward games, deploy-time certificates: the fixes that moved agent quality this week did not...

Hex's data-agent lab shows the practical eval pattern AI teams should copy: compare candidates against stable baselines,...

A new 188-task benchmark for non-functional improvements finds coding agents hit 70% on functional correctness but lag h...

The first production-scale trace of agentic coding says the context you keep paying for is already dead at every turn bo...

New tutorials, open-source projects, and deep dives on coding agents - delivered weekly.