LLM Judge Validation Under Sparse Overlap: From Inference to Design

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of erroneous deployment decisions in LLM-as-a-judge validation, which arise from sparse sample overlap due to constrained annotation budgets. We present the first quantitative analysis establishing the decisive impact of overlap sparsity on decision errors, derive a closed-form expression for the minimum required overlap, and propose a zero-cost stratified sampling allocation strategy grounded in statistical inference. Extensive experiments across ten LLM judges demonstrate that the proposed approach reduces the false rejection rate by 50% and identifies ρ ≥ 0.25 as a critical safety threshold for reliable evaluation. By bridging theoretical guarantees with practical efficiency, this work provides a principled framework for the robust and cost-effective deployment of large language models under limited annotation resources.
📝 Abstract
Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this \emph{overlap sparsity} is the first-order determinant of wrong deployment decisions: at 5\% pairwise overlap, wrong-decision rates reach 25\% and the probability of selecting the wrong best judge among ten candidates is 65\%. The two actionable levers are overlap \emph{quantity} and \emph{allocation}. For quantity, we derive a minimum-overlap formula showing $\rho \geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-a-judge
judge validation
sparse overlap
annotation budget
human agreement
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-as-a-judge
overlap sparsity
minimum-overlap formula
stratified allocation
validation design
🔎 Similar Papers
No similar papers found.