Agent / Code / Safety
AgentGuard: Learning Execution Guardrails from Anomalous Coding-Agent Trajectories
Wuyang Dai, Song Wang
AgentGuard: Learning Execution Guardrails from Anomalous Coding-Agent Trajectories
Authors: Wuyang Dai, Song Wang
arXiv ID: 2609.16287
Problem: Coding agents succeed at tasks while executing unreliably - modifying unrelated files, rewriting tests, issuing unsafe commands, ignoring failed validations. Hand-written safety rules do not scale and over-restrict normal execution; the alternative is to learn behavioral guardrails from what agents actually do wrong.
Key Methodology:
- 642 documented failure traces collected from real coding-agent executions across 382 repository tasks
- Recurring execution failure patterns are extracted and generalized into conditional instruction-level behavioral constraints, organized as a lightweight guardrail skill that dynamically activates only the rules relevant to the current instruction
- Guardrails learned from 461 traces covering 282 tasks; evaluated on a disjoint set of 100 tasks
- Baseline versus augmented agent: Claude Code with Claude Haiku 4.5 as the underlying coding agent
Key Results:
- Abnormal Execution Rate reduced from 69.0% to 26.7%
- Successful Task Completion Rate increased from 21.7% to 35.0%
- Instruction-conditional activation keeps restrictions minimal: only rules relevant to the current instruction fire
Applied Context: The failure corpus of an agent fleet is a training signal: mined, conditional guardrails beat both unscalable hand-written rules and blanket restriction on a fresh task set. The residual 26.7% abnormal rate is the measured reminder that learned guardrails narrow the safety/completion tradeoff but do not close it.
Paper: arXiv:2609.16287