JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational cost of employing large language models (LLMs) as evaluation judges by investigating the substitution potential of a lightweight classifier, JEV. Across nine benchmarks, JEV is compared against three LLMs, incorporating typed classifiers, cross-validated thresholds, and cost-benefit analyses to systematically evaluate accuracy, overhead, and error correlation, while exploring cascaded application strategies. Results demonstrate that JEV achieves 29–325× cost reduction and 30–220× speedup relative to LLMs. However, because errors between JEV and LLMs are highly correlated, the cascading strategy yields only a marginal 1.5-point improvement over the best single judge. This work reveals the efficiency advantages of lightweight judges while exposing the fundamental limitations in performance gains achievable through cascading architectures.
📝 Abstract
We ask whether Jev, a typed classifier that returns probabilities over permitted answers without generating text, can replace an LLM rubric judge. We compare it with three flash-tier LLM judges on nine panels drawn from seven benchmarks, giving every judge identical criterion texts. Jev's accuracy differs significantly from an LLM judge's in only 8 of 27 paired comparisons, ahead mostly on binary criteria and behind only on graded ones, and most of the other comparisons are inconclusive. Summed over the nine panels, the LLM judges, called once per criterion, cost 29 to 325 times as much as Jev and took 30 to 220 times as long. On graded criteria all four judges agree more with one another than with the labels and mostly assign lower levels than the raters. One of several observational accounts is that raters followed scale conventions our criterion texts omit. Jev's confidence ranks its own errors on most panels, which should make a cheap classifier the ideal first stage of a cascade that defers its uncertain verdicts to an LLM judge. Correlated errors undo that advantage. The LLM judges repeat nearly all of Jev's most confident errors, so a cascade replayed on the recorded verdicts lowers cost but gains at most 1.5 points over the best single judge with cross-fitted thresholds, and at most 2.0 even with oracle thresholds.
Problem

Research questions and friction points this paper is trying to address.

Rubric Judge
LLM Evaluation
Typed Classifier
Cascade System
Graded Criteria
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rubric Judge
Typed Classifier
Cascade System
Correlated Errors
LLM Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.