🤖 AI Summary
This study addresses systematic deficiencies in the LLM-as-a-Judge paradigm, including positional bias, irreproducibility, and misalignment with human judgment. Through multi-benchmark stress testing and controlled experiments, we comprehensively audit the reliability of judgments produced by frontier models. We propose a "trustworthy judgment rate" metric to quantify evaluation stability and derive the theoretical upper bound on accuracy imposed by positional bias. Our findings reveal that judgments remain unstable even at zero temperature, and that reliability is sample-specific rather than an intrinsic model-level property. Furthermore, transitioning from pairwise comparison to holistic scoring yields substantially greater improvements in trustworthiness than single-prompt interventions. This work provides both a theoretical foundation and a practical framework for constructing robust NLP evaluation systems.
📝 Abstract
LLM-as-a-Judge is the standard paradigm for NLP evaluation, yet its systemic reliability remains poorly understood despite being widely treated as a deterministic ground truth. We present a comprehensive reliability audit, stresstesting six frontier models across four benchmarks, five prompt formats, two presentation orders, three sampling temperatures, and ten repetitions per condition. Our empirical analysis reveals severe vulnerabilities: verdicts change across identical replications at temperature zero, position-order swaps flip the majority of verdicts on challenging tasks, and the most deterministic judge achieves perfect consistency by trivially repeating incorrect verdicts, agreeing with ground truth only 51% of the time. To formalize these multi-faceted failure modes, we introduce the trustworthy verdict rate (T ), a unified metric capturing the joint probability that an evaluation is reproducible, order-invariant, and accurate. UsingT , we derive a theoretical upper bound on accuracy imposed by position bias and show that reliability is item-specific rather than modellevel. Finally, we demonstrate that shifting from pairwise win-rate to holistic rubric scoring improves trustworthiness more than any single-format prompting intervention, offering an actionable framework for robust NLP evaluation.