Agents / Evaluation / Cost
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman · Carnegie Mellon University
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
Authors: Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman
arXiv ID: 2609.26550
Problem: LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. Can a decision-only judge provide an economical first pass, and identify when stronger evaluation is actually needed?
Key Methodology:
- A decision-only judge (JEV) compared against sixteen generative and reward-model judges
- Blinded human adjudication as the reference
- Evaluated on ordinary preference judgments and evidence-grounded factuality, plus harder cases: checking a derivation and resisting an elaborately written wrong answer
- Frozen cascade design: accept confident verdicts, escalate uncertain ones to the stronger judge
Key Results:
- JEV is within three percentage points of the strongest comparator (a state-of-the-art LLM judge) on ordinary preference and evidence-grounded factuality, at 0.36% of the comparator's fee
- Larger gaps arise when judgments require checking a derivation or resisting an elaborately written wrong answer
- On several benchmarks the gap concentrates in JEV's low-confidence decisions
- The frozen accept/escalate cascade retains 99% of the comparator's accuracy at lower cost
What it means for developers: Verdict cost collapses to 0.36% of an LLM-judge pass for the large majority of ordinary judgments, with the expensive judge spent only on the uncertain tail - the confidence signal does the routing. Before budgeting frontier-judge tokens per decision in an eval loop, measure what fraction of your judgments a decision-only head can clear; a frozen cascade is now a measured architecture, not a hope.
Paper: arXiv:2609.26550