Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear statistical validity of randomization and averaging strategies in large language model (LLM) evaluation, where nonlinearity often yields inconsistent conclusions. We establish the first theoretical framework based on Markov generalized linear mixed models (GLMMs) to approximate LLM evaluation mechanisms. Methodologically, by integrating sequential effect modeling with statistical inference techniques, we reveal the inconsistency of naive averaging and propose a Williams design to enhance evaluation efficiency. Experiments conducted on three commercial LLMs validate the predictive capability of our model, demonstrating that GLMM-based statistical inference remains effective even under higher-order sequential memory. This work provides an optimized experimental design paradigm for leaderboard ranking and inter-group comparisons in LLM evaluation.
📝 Abstract
Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear mixed models (GLMMs), supported by out-of-sample predictions across three major commercial LLMs. Using a first-order Markov GLMM, we study leaderboard ranking and group comparison. For leaderboard ranking, randomize-and-average selection is consistent under a mild separation condition, and a Williams square design can improve efficiency when item qualities are close. For group comparison, naive averaging can yield inconsistent conclusions about differences in group-level quality because of the response model's nonlinearity. Empirical results further support the validity of the proposed model-based inference beyond the first-order theory, including settings with higher-order sequence memory. We illustrate the approach in an application where AI judges compare two graphical model estimation methods.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-a-judge
method comparison
statistical validity
order effects
aggregation effects
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-as-a-Judge
Markov GLMM
Williams square design
leaderboard ranking
model-based inference
🔎 Similar Papers
No similar papers found.