🤖 AI Summary
This study addresses the high computational cost of employing large language models (LLMs) as evaluation judges by investigating the substitution potential of a lightweight classifier, JEV. Across nine benchmarks, JEV is compared against three LLMs, incorporating typed classifiers, cross-validated thresholds, and cost-benefit analyses to systematically evaluate accuracy, overhead, and error correlation, while exploring cascaded application strategies. Results demonstrate that JEV achieves 29–325× cost reduction and 30–220× speedup relative to LLMs. However, because errors between JEV and LLMs are highly correlated, the cascading strategy yields only a marginal 1.5-point improvement over the best single judge. This work reveals the efficiency advantages of lightweight judges while exposing the fundamental limitations in performance gains achievable through cascading architectures.
📝 Abstract
We ask whether Jev, a typed classifier that returns probabilities over permitted answers without generating text, can replace an LLM rubric judge. We compare it with three flash-tier LLM judges on nine panels drawn from seven benchmarks, giving every judge identical criterion texts. Jev's accuracy differs significantly from an LLM judge's in only 8 of 27 paired comparisons, ahead mostly on binary criteria and behind only on graded ones, and most of the other comparisons are inconclusive. Summed over the nine panels, the LLM judges, called once per criterion, cost 29 to 325 times as much as Jev and took 30 to 220 times as long. On graded criteria all four judges agree more with one another than with the labels and mostly assign lower levels than the raters. One of several observational accounts is that raters followed scale conventions our criterion texts omit. Jev's confidence ranks its own errors on most panels, which should make a cheap classifier the ideal first stage of a cascade that defers its uncertain verdicts to an LLM judge. Correlated errors undo that advantage. The LLM judges repeat nearly all of Jev's most confident errors, so a cascade replayed on the recorded verdicts lowers cost but gains at most 1.5 points over the best single judge with cross-fitted thresholds, and at most 2.0 even with oracle thresholds.