Skip to main content

Agents / Evaluation / Cost

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman · Carnegie Mellon University

arXiv:2609.2655012 upvotes

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Authors: Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman

arXiv ID: 2609.26550

Problem: LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. Can a decision-only judge provide an economical first pass, and identify when stronger evaluation is actually needed?

Key Methodology:

  • A decision-only judge (JEV) compared against sixteen generative and reward-model judges
  • Blinded human adjudication as the reference
  • Evaluated on ordinary preference judgments and evidence-grounded factuality, plus harder cases: checking a derivation and resisting an elaborately written wrong answer
  • Frozen cascade design: accept confident verdicts, escalate uncertain ones to the stronger judge

Key Results:

  • JEV is within three percentage points of the strongest comparator (a state-of-the-art LLM judge) on ordinary preference and evidence-grounded factuality, at 0.36% of the comparator's fee
  • Larger gaps arise when judgments require checking a derivation or resisting an elaborately written wrong answer
  • On several benchmarks the gap concentrates in JEV's low-confidence decisions
  • The frozen accept/escalate cascade retains 99% of the comparator's accuracy at lower cost

What it means for developers: Verdict cost collapses to 0.36% of an LLM-judge pass for the large majority of ordinary judgments, with the expensive judge spent only on the uncertain tail - the confidence signal does the routing. Before budgeting frontier-judge tokens per decision in an eval loop, measure what fraction of your judgments a decision-only head can clear; a frozen cascade is now a measured architecture, not a hope.

Paper: arXiv:2609.26550