Anytime-Valid LLM Leaderboards via Benchmark-weighted and Block-Factorized e-Processes

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of unwarranted superiority claims on LLM leaderboards caused by evaluation volatility and frequent updates. To this end, it proposes the BB-EDGE framework, which models leaderboards as directed graphs and introduces a novel protocol-defined block-based evidence factorization method. By integrating benchmark weighting with block-decomposed e-processes, the framework achieves anytime-valid family-wise error rate (FWER) control under arbitrary dependency structures. Specifically, the method employs empirical Bernstein e-processes, direct e-Holm correction, and weight-proportional linear betting strategies to support top-k certification and simultaneous confidence interval construction. Experiments on synthetic data and four real-world benchmarks validate its rigorous error control capabilities and statistical efficiency.
📝 Abstract
Large language model (LLM) leaderboards compare model capabilities by ranking models according to their mean performance on fixed benchmarks. However, variability in evaluation outcomes across runs may produce unsupported claims of model superiority on the benchmark, a risk compounded by leaderboard updates. In this paper, we propose BB-EDGE (Benchmark-Weighted and Block-Factorized e-processes for Directed Graph Evaluation), a principled framework that represents an LLM leaderboard as a directed graph whose edges certify pairwise mean-performance advantages, with anytime-valid family-wise error rate (FWER) control. Concretely, for each direction, BB-EDGE constructs an empirical-Bernstein e-process by factorizing evidence over protocol-defined blocks and assigning stakes proportional to the corresponding block weights, then applies direct e-Holm across these $e$-processes to certify directional advantages as edges. Theoretically, we characterize weight-proportional linear stakes under heterogeneous benchmark-average nulls and prove anytime FWER control under arbitrary within-block and cross-pair dependence. BB-EDGE further supports anytime-valid Top-$k$ certification and simultaneous rank intervals. Extensive experiments on synthetic data and four real-world benchmarks demonstrate that BB-EDGE maintains anytime FWER control while achieving high efficiency.
Problem

Research questions and friction points this paper is trying to address.

LLM leaderboards
evaluation variability
anytime-valid inference
family-wise error rate
model comparison
Innovation

Methods, ideas, or system contributions that make the work stand out.

Anytime-valid inference
e-processes
Family-wise error rate control
LLM leaderboards
Directed graph evaluation
🔎 Similar Papers
No similar papers found.