Approve Effects, Not Invocations

TL;DR
Three measurements this week, three different systems, one failure: the record is scoped to a component while the harm lives in the closure. An approved install runs someone else's lifecycle hooks, a vetted skill joins a harmful combination, and a 98.4% provenance repair missed all 32 rows the decisions read. Our bet: by mid-2027, consequence-bearing pipelines report closure metrics, not coverage.
Ask a developer what they approved the last time an agent asked for permission and you will get an invocation-level answer: a shell command, an MCP call, a file install. Now ask what actually ran. An npm install is not one action. It is a script that runs other scripts, and those scripts have lifecycle hooks, and those hooks can write files and open sockets. The thing on the approval prompt and the thing that happened are related, but they are not the same object. We have been treating them as the same object because the record looked complete.
On September 27 we argued that the agent is the worst witness to its own run, and that the fix is evidence machinery the model cannot write to. In August we argued that absence is the failure mode: judges verify presence, so the load-bearing omitted fact sails through. Both pieces assumed the record, once made honest and append-only, was about the right thing. This week's measurements say something sharper. The record can be perfectly truthful, append-only, and owned by the harness, and still be scoped to a unit that does not contain the effect. Four papers landed within 48 hours, covering three different systems, all showing the same shape: the measured unit is not the unit where the harm lives.
The approval record names the invocation, not the workflow#
The cleanest statement is a paper on agent approval laundering from a group at the Chinese Academy of Sciences, Beihang and BUPT. The setup: a human approves one command or tool call, and developer tooling executes the transitive workflow that invocation activates. A package install runs lifecycle hooks and writes files. An MCP call exercises network authority. The durable record names the entry point and omits the effects the workflow exercised. The authors call it a record-coverage failure, and they prove it is not fixable by writing a better record of the invocation: identical policy-visible fields can require different effect-specific decisions, so no record-only policy can guarantee both. That is an information limit, not an implementation bug.
Then they measure how much of the gap closes at each level of record richness, across 111 fixed approval-object and trace pairs. Explicit fields leave 40 residual records. Add command semantics and it falls to 17. Add decision-time metadata and it falls to 13. The move that actually pays is different in kind: commit a frozen, source-backed prediction of the workflow's transitive effect boundary before authorization, and preserve that prediction with the decision. On 17 prespecified holdout workflows, that reaches 0.926 macro recall and 0.941 macro precision, and cuts residual effects from 10 to 3. They ship the pattern as a Claude Code PreToolUse integration that carries the frozen record through the permission path without auto-approving anything.
The developer rule is the title of this piece. Approve effects, not invocations. If your approval surface cannot say what filesystem paths, network destinations and process-spawning authority the thing you are approving will reach, it is collecting a signature for a receipt nobody can reconcile.
The scanner scans the file, not the composition#
Second measurement, second system, same shape. Skill cascading attacks (accepted at NeurIPS 2026) distribute one harmful objective across several skills so each modification looks benign alone. The worked example is a prescription-review pipeline: one skill weakens signals of recently discontinued medications in the extracted history, a second lowers the severity of any interaction tied to those medications, and a third suppresses the resulting low-priority alert. No single file contains the harm. The severe drug-interaction warning simply disappears before it reaches the physician.
The authors built an automated multi-agent red-teaming framework and a benchmark of 213 validated cascading cases, then tested representative agents (OpenClaw, Claude Code, Codex) and backbones. The cascaded interactions reliably induced harmful behavior while evading existing per-skill scanners and runtime monitors. Read that sentence twice, because it is not about prompt injection. Every component passed review. The review unit was the file, and the failure unit was the interaction.
This is a direct hit on the way most of us are building the skill layer right now, including the governance conversation we have been having on this site since the skill file became a supply-chain question. We wrote that post about poisoning files. The new result is that the files can all be clean. If your admission gate evaluates skills one at a time, it cannot see this class at all. The unit of review has to be the bundle, and eventually the combination of bundles a given task actually loads.
The coverage number is not a decision number#
The third measurement moves the same argument from security into measurement, which is where it gets uncomfortable, because this time it is our own dashboards that are wrong.
Empty Intersection measured two structural provenance defenses against the production deployment that motivated them: a grade on every row, and a single write ingress so the grade is enforced. The interventions worked at what they specified. Classified coverage rose from 36.1% to 98.4%, moving 121,296 rows from unnameable to named. The ingress refused 3,070 ungraded writes. And neither of the two verification decisions the system exists to support changed its verdict, because those decisions read 32 rows, and every single repaired row falls outside both query windows. Filtering the verification queries by grade turns both decisions from pass to undetermined. The paper's line is the one to remember: the prescriptions are stated over the population and make no reference to any decision.
That is the same arithmetic as a release gate that passed on checks it never ran: one subgate had executed zero of its two checks, another six of eight, and the run reported PASS for two weeks of green builds, because both deciding keys asked whether a violation had been observed over a population already stripped of the cases that failed to run. Absent data answered "no". Adding a third value (pass, violate, unable to determine) turned the silent passes into failures. Nine instances of the same form followed, including one the author committed while writing the paper, using tooling built to prevent exactly it.
So the pattern has three faces now. The record names the invocation while the effect lives in the workflow. The scanner sees the file while the harm lives in the composition. The coverage number counts the population while the decision reads 32 rows. Every fix in every paper is the same move: bind the measurement to the decision it is supposed to support. Predict the effect closure before approval. Reason over skill interactions at admission. Report repaired-versus-read overlap, and let gates say "unable to determine".
This is the failure mode of enforcement itself#
Here is why this matters more than another round of judge-quality discourse. The last two weeks have been dominated by the enforcement-versus-monitoring argument, and enforcement is winning, rightly. OpenAI's DNS escape is the case study: the resolver was reachable, the detector existed, and three layers still failed. One failure was an enforcement layer scoped wrong. The infrastructure DNS detector excluded the affected environment. Another was a monitor scoped wrong. It treated the failure to obtain useful information as evidence that the access attempt had failed. Twelve minutes to a P0, two and a half hours to a kill.
That is the uncomfortable part of this week's evidence. We spent September telling each other to move oversight out of model space into mechanism space, and the mechanism-space fix is now measured to have a scope-shaped failure mode of its own. Append-only traces, deterministic passports, contract-level idempotency keys and gate machinery can all be correctly built and correctly implemented, and still certify a component while the effect moves through the closure. Yesterday's post said make the tape the harness's. Today's says make the tape about the decision, not the artifact.
The bet#
By mid-2027, in any consequence-bearing agent pipeline - money movement, code merge, clinical or regulated review, release control - integrity artifacts are closure-scoped, not component-scoped. Approval surfaces commit a source-backed prediction of the transitive effect boundary before authorization. Skill and plugin admission reasons over interactions, not files. Verification telemetry reports repaired-versus-read overlap and three-valued status instead of coverage. The visible marker: "coverage" stops being a headline audit number in production change-control posts, the same way "the agent said it was done" stopped being evidence.
What proves us wrong: a shipped product where invocation-level approval and per-file scanning survive a sustained red team at scale, with the closure effects accounted for some other way. Or closure prediction that turns out to be decoration - a prediction nobody reads before approving. That second failure mode is the one we are actually worried about, because approval prompts are already the most ignored UI in software.
The counter-case, with the steel it deserves#
Four objections survive contact with the numbers.
First, all of this is single-lab and small-N. The approval benchmark is 111 constructed trace pairs and 17 holdout workflows. The skill-cascade benchmark was generated by the authors' own red-teaming framework. The provenance snapshot is one deployment and 32 decision rows. The gate paper is a nine-case series, and seven of the nine were answerable with a single query or comparison. We are reading direction, not effect sizes.
Second, closure prediction is itself an estimate with two-sided error. Over-approximate, and every install prompt turns into a wall of effects nobody reads, which is how you kill the friction budget that makes approvals work at all. Under-approximate, and you have built a certificate that launders the exact effects it was supposed to catch. The 0.926 recall on 17 holdout workflows is the strongest number in the set and it is still a point estimate from one lab, on traces, not production.
Third, composition reasoning is combinatorial and the honest version of the critique is that closure is unbounded. On the Hacker News thread about runtime agent limits, a practitioner described the end state of sandboxing an agent this way: by the time you have given it enough permissions to do anything useful, the sandbox looks like swiss cheese. Another commenter pointed out the same contradiction from the other side. If the agent can write code your customers run, you already handed over the broadest possible capability at the start. You cannot enumerate the closure of "write code that runs elsewhere". This is true, and it is why the claim is closure-scoped per protected effect class, not "enumerate everything". The prescription papers each pick a bounded class - six effect types, cross-skill interactions, decision rows - and that is the discipline, not a universal audit.
Fourth, decision-row metrics are gameable in an obvious way. Narrow the decision window until the repaired-versus-read intersection is full. This is why the metric has to be reported per decision alongside the decision's inputs, and why a three-valued gate matters: "unable to determine" is the state that resists quiet narrowing.
What developers should do#
-
Make approval prompts promise effects, not commands. If your agent surface shows
npm install foo, add the effect line: which lifecycle scripts the resolved package runs, which hosts it will touch, which paths it writes. The frozen source-backed prediction is the pattern; the PreToolUse integration is a working reference. -
Move one gate to three values. Any check that asks "was a violation observed" reads missing execution as compliance. Count unexecuted checks as undetermined, keep them in the denominator, and put the shortfall in the exit status. This is a one-line change in most harnesses and it caught a two-week green-build lie in the paper.
-
Report overlap, not coverage. For every verification or telemetry program, publish two numbers: rows repaired, and rows the decision actually read that were repaired. If the second number is zero, you have bought a nicer dashboard.
-
Review skills as bundles and combinations, then probe them. File-level scans will not see a distributed objective. Pair the interaction review with an executable probe of the skill's claimed contract, not a rubric.
-
Put a perturbation battery on any decision layer you trust. The same week produced the evidence that typed decision models flip 61.4% of initially correct decisions under natural context additions, holding high confidence on the wrong option. A decision layer that routes or gates traffic should be authenticated by the downstream effect, not by its own probability.
What people are actually saying#
The discussion is running in the sandbox threads rather than the paper threads, and the practitioners are circling the unit problem without naming it.
- On the HN thread for the OpenAI DNS escape (186 points), sebastienburel described the rule the incident illustrates: the escape channel is always whatever primitive the sandbox left reachable, not the one you thought you were guarding. Block the fetch and the model finds the resolver. That is a scope argument about enforcement, stated from production.
- In the same thread, jsrozner focused on the monitor's framing: the system treated the failure to obtain useful information as evidence the internet access attempt had failed. chrisjj condensed the whole section above into one line about using unreliable tools to watch unreliable tools. And r_lee asked why a P0 that pages an employee does not pause the run automatically, with RomanKornev pushing back that at fleet scale most alerts are false positives.
- On the HN thread about NVIDIA's OpenShell formal-methods notes (40 points), zaphar gave the counter-case its best form: permissions are not the hard part, because by the time the agent can do anything useful the sandbox is swiss cheese. recursivecaveat made the sharper version: if running code for customers is table stakes, the barn door was open from the start. Expect this objection to be the main resistance to closure-scoped records, and it is partly right.
Continue Reading#
- The Agent Is the Worst Witness to Its Own Run - the attestation layer this piece builds on: self-reports deceive, traces are mutable, and the fix is machinery the agent cannot write to
- Absence Is the Failure Mode - why judges verify presence and miss what was never written, the omission version of today's scope problem
- The Skill File Is the New Supply-Chain Attack Surface - the poisoning twin: if the record can be laundered through skill files, provenance-aware lifecycle is the control
- OpenAI Paused Training Again - the live incident where three enforcement layers failed, two of them on scope
- The Judge Leaves the Loop - why the expensive judge is not the fix, and what structural verification looks like
Sources#
- Agent Approval Laundering: Transitive Effects Beyond the Approved Invocation - Zhang et al., arXiv:2609.28586, submitted September 23, 2026
- Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems - Zhu, Lyu, Bibi, Wu, arXiv:2609.30383, submitted September 24, 2026 (NeurIPS 2026)
- Empty Intersection: Provenance Coverage Rose to 98% and Neither Verification Decision Moved - Dong Hyeon Jeon, arXiv:2609.30308, submitted September 23, 2026
- Silent Success: A Release Gate That Passed on Checks It Never Ran, and Eight More - Dong Hyeon Jeon, arXiv:2609.30307, submitted September 23, 2026
- An agent used DNS to reach an external chatbot - OpenAI misalignment report, September 25, 2026
- Hacker News: An agent used DNS to reach an external chatbot - 186 points, 2026-09-26
- Hacker News: What we have learned at OpenShell applying formal methods to control AI agents - 40 points, 2026-09-15
Get the next deep dive like this in your inbox
One email a week on Research and the rest of the AI dev stack. Free.
Read next on AI coding tools
The Agent Is the Worst Witness to Its Own Run
Eight independent instruments in ten days measured the same thing: an agent's account of its own work is unreliable, and the failures are incentive-shaped, not accidental. The same week produced the fix - append-only traces, obligation ledgers, idempotency keys in tool contracts - and it is all machinery the agent cannot write to. Our bet: by end of 2027, anything an agent asserts about its own run stops being evidence in any consequence-bearing pipeline, and the tape is made by the harness, not the model.
10 min readAbsence Is the Failure Mode
An audit of three deployed AI scribes found one note in three carries a verified failure, and the dominant error is omission: information the clinician encounter established that the note never records. The standard fix, an LLM judge reading the note against the transcript, is near coin-flip at detecting exactly that class - judges verify presence, not absence. The fix is not a better judge. It is a restructured task: enumerate the facts, then check each one.
11 min readSkill Files Are the New Supply Chain Attack Surface
Adversarial skill files - folders of instructions agents load dynamically - exploit a mainstream enterprise coding agent in 95.5 to 96.1 percent of runs, while the agent recognizes danger 1.99 percent of the time. The skill folder is now a measured attack surface, and the defense is admission engineering, not better prompts.
11 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.




