Agents / SRE / Evaluation
Correct Source, Empty Story: Why Microservice RCA Agents Need Trajectory-Level Grading
Qisheng Lu, Aoyang Fang, Junjielong Xu, Jin'ao Shang, Songhan Zhang, Yifan Yang, Xiaochuan Yan, Pinjia He
Correct Source, Empty Story: Why Microservice RCA Agents Need Trajectory-Level Grading
Authors: Qisheng Lu, Aoyang Fang, Junjielong Xu, Jin'ao Shang, Songhan Zhang, Yifan Yang, Xiaochuan Yan, Pinjia He
arXiv ID: 2608.21310
Problem: Automated root cause analysis (RCA) for microservices is almost always graded on endpoint correctness: did the agent name the responsible service. That criterion enables apples-to-apples comparison but ignores what an on-call SRE actually needs before they act - the evidentiary basis of the diagnosis and the fault-propagation route connecting source to symptom. No existing evaluation captures whether the agent can reconstruct how the failure spread.
Approach: Treat RCA as an observable diagnostic process. The authors build a trajectory-level framework that grades agent executions against manually curated, service-level fault-propagation paths on a public microservice RCA benchmark, then analyzes 3,500 diagnostic trajectories to characterize where agents investigate and how they use retrieved telemetry. From the failures they derive a taxonomy and operationalize it as DiagGuard, a two-stage defense where a grounding stage surveys available observations before the localization stage.
Key Results:
- Answer correctness and diagnostic quality are disconnected: an agent can localize the fault source yet fail to reconstruct its propagation
- Successful investigations stay on the fault-impact surface, act on retrieved evidence, and broaden their query repertoire as the search deepens
- Failures fall into three classes: decisive evidence omitted, retrieved evidence misinterpreted, and unsupported inference substituting for missing evidence
What it means for developers: If you evaluate an oncall or RCA agent only by "did it name the right service", you will ship agents that pass the check while telling an empty or wrong story - and an SRE cannot act on a name alone. Grade the evidentiary chain: does the agent ground its claim in observed telemetry, and does it reconstruct the propagation path, not just the source? For teams running agent triage, DiagGuard's grounding-before-localization shape is a deployable harness pattern, and it extends the case that today's agent RCA is not yet trustworthy enough to act on unattended (see our ORCA-bench coverage).
Paper: arXiv:2608.21310