How to Run Statistics over LLM Judges and Trust the Results: Calibrated Inference for Small-Sample AI Evaluation with evalstats

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unreliability of statistical inference for LLM-as-a-judge evaluations in small-sample settings, where inflated false positive rates are particularly pronounced under high inter-rater agreement. To mitigate this, the authors propose a calibration framework grounded in Prediction-Powered Inference (PPI), introducing novel PPI correction methods tailored for four rank-based tests. By integrating Monte Carlo simulations with bootstrap techniques, the approach enables adaptive power tuning. Furthermore, the work presents evalstats, an open-source Python toolkit that automatically selects optimal calibration strategies to prevent spurious discoveries. This contribution significantly enhances the statistical reliability of few-shot evaluations, providing rigorous hypothesis testing support for LLM-based assessment paradigms.
📝 Abstract
Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated confidence intervals (CIs), hypothesis tests, and judge-bias corrections, such claims are unreliable. We address these issues in several contributions. First, we find that running statistics over raw LLM judge scores leads to inflated false positives: counterintuitively, for many inter-rater agreement metrics, false positive risk peaks at "almost perfect" human-LLM agreement. To help researchers understand how to run statistics over LLM judges responsibly, we present guidance and tooling for the statistical analysis of mixed human-AI judge designs, and implement nine hypothesis tests via prediction-powered inference (PPI), including the first known PPI corrections for four rank-based tests (Wilcoxon signed-rank, Mann-Whitney U, and omnibus variants). To keep PPI++ stable with small human-labeled calibration sets, we introduce bootstrap-adaptive power tuning, which shrinks the estimated weight toward a target estimated from the labeled data, and accounts for that weight's own sampling variance. Second, through Monte Carlo simulations, we derive recommendations for what CI, p-value, and FWER correction methods to use for small-sample AI evaluations (N<100), and warn researchers against bootstrap CIs. We package these recommendations into evalstats, an open-source Python package that selects calibrated methods automatically, and demonstrate it in three scenarios, including one where a real LLM judge validated at "substantial agreement" would have led a researcher to publish a spurious finding. evalstats is publicly available at https://github.com/ianarawjo/evalstats.
Problem

Research questions and friction points this paper is trying to address.

LLM judges
small-sample AI evaluation
statistical significance
false positives
confidence intervals
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prediction-Powered Inference
LLM Judges
Small-Sample Evaluation
Bootstrap-Adaptive Power Tuning
evalstats
🔎 Similar Papers
No similar papers found.