JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过JEV-as-a-Judge方法在保证低成本的同时,对任务进行初步评估并识别需要更强评估的情况,以解决大规模应用中推理成本和信心可靠性的问题。
📝 Abstract
LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking a derivation or resisting an elaborately written wrong answer. On several benchmarks, JEV's gap to this comparator is concentrated in low-confidence decisions. A frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% of the comparator's accuracy at lower cost.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-a-judge
inference cost
confidence reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

JEV-as-a-Judge
Decision-Only Judge
Cost Efficiency
Confidence-Based Escalation
Frozen Cascade