🤖 AI Summary
This study addresses the challenge of erroneous deployment decisions in LLM-as-a-judge validation, which arise from sparse sample overlap due to constrained annotation budgets. We present the first quantitative analysis establishing the decisive impact of overlap sparsity on decision errors, derive a closed-form expression for the minimum required overlap, and propose a zero-cost stratified sampling allocation strategy grounded in statistical inference. Extensive experiments across ten LLM judges demonstrate that the proposed approach reduces the false rejection rate by 50% and identifies ρ ≥ 0.25 as a critical safety threshold for reliable evaluation. By bridging theoretical guarantees with practical efficiency, this work provides a principled framework for the robust and cost-effective deployment of large language models under limited annotation resources.
📝 Abstract
Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this \emph{overlap sparsity} is the first-order determinant of wrong deployment decisions: at 5\% pairwise overlap, wrong-decision rates reach 25\% and the probability of selecting the wrong best judge among ten candidates is 65\%. The two actionable levers are overlap \emph{quantity} and \emph{allocation}. For quantity, we derive a minimum-overlap formula showing $\rho \geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.