statistical leaderboard analysis

Designs and implements methods and tools to manage leaderboards and perform statistically rigorous evaluation of models and benchmarks. This includes computing per-task bootstrap confidence intervals, running paired-bootstrap significance tests, applying item-response-theory discrimination analyses, and measuring cross-leaderboard correlations to enable robust model comparison and uncertainty-aware ranking.

statisticalleaderboardanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a critical limitation in current model leaderboard evaluation methods, which often overlook task-level performance uncertainty and variability, thereby yielding unreliable global rankings. To remedy this, the authors propose a hierarchical framework that enables principled aggregation of uncertainty from the task level to the leaderboard level. By integrating task-wise pairwise comparisons with conformal prediction, the method produces statistically valid confidence intervals for both individual task rankings and overall model rankings. Empirical validation on synthetic data as well as real-world benchmarks—TabArena and PromptEval (MMLU)—demonstrates the approach’s statistical validity and informativeness. The resulting rankings are not only reliable but also explicitly account for uncertainty, offering a robust foundation for evaluating and comparing models on new tasks.

leaderboardmodel rankingmulti-task evaluation

This study addresses the problem of uncontrolled confidence interval error rates in dynamic evaluation of model leaderboards caused by repeated testing or early stopping. We propose an anytime-valid confidence sequence method for full-model ranking that combines betting e-processes for pairwise comparisons with closed testing to integrate all possible orderings. This approach constructs, for the first time under score-dependent conditions, confidence sequences that permit arbitrary stopping rules without inflating error rates. Simulations and experiments on public datasets demonstrate that the proposed method supports real-time monitoring and adaptive stopping, significantly reducing computational costs while sacrificing only minimal statistical power compared to fixed-sample approaches.

anytime-valid leaderboardsdependent scoreserror rate control

Stratified Bootstrap Test Package

Dec 16, 2025
EM
Ehsan Mohammadi

This study addresses the stability assessment of group-specific ranking patterns and nonparametric inference on population-mean ordinal relationships in multivariate survey/scoring data. We propose the first hierarchical bootstrap framework for ordinal hypothesis testing, which approximates the null distribution without distributional assumptions. We introduce the *non-containment index*—a robust, interpretable measure quantifying ranking stability across groups—and leverage it for outlier response detection and significance testing of inter-group ranking differences. The method integrates hierarchical resampling, nonparametric stability analysis, and resampling-based ordinal inference, unifying descriptive and inferential capabilities. Evaluated in AI fairness auditing and questionnaire analysis, it demonstrates high sensitivity and reliability. Our approach establishes a novel, assumption-free, robust, and interpretable statistical paradigm for ordinal data analysis.

Assesses stability of group-specific ranking patterns in multivariate dataEnables descriptive and inferential evaluation of ranking consistency across groupsQuantifies ranking robustness using a non-containment index via resampling

Bootstrapping the Cross-Validation Estimate

Jul 01, 2023
BC
Bryan Cai
🏛️ Stanford University | Biogen Inc

This paper addresses the challenge of quantifying uncertainty in cross-validation (CV) performance estimates—particularly the difficulty of distinguishing true performance differences from random fluctuations during model comparison. We propose an efficient and robust bootstrap-based method that decomposes the variance of CV estimates using a random-effects model, enabling valid statistical inference on CV performance differences without strong modeling assumptions. Compared to standard bootstrap, our approach substantially reduces computational cost while yielding confidence intervals with accurate coverage probability and high statistical power. Extensive evaluations—including simulations and real-world applications across classification, regression, and time-series tasks—demonstrate its strong robustness and generalizability under small-sample settings, non-independent CV folds, and heterogeneous data distributions. The method provides reliable uncertainty quantification to support principled model selection.

Comparing model performance differences without stringent assumptionsEstimating uncertainty in cross-validation performance estimatesProviding computationally efficient confidence intervals for prediction models

Quantifying Uncertainty: All We Need is the Bootstrap?

Mar 29, 2024
UZ
Urvsa Zrimvsek
🏛️ University of Ljubljana

This paper addresses the complexity and high pedagogical/practical barriers associated with conventional uncertainty quantification methods—such as standard errors, confidence intervals, and hypothesis tests—in statistical inference. To evaluate the potential of nonparametric bootstrap as a unified alternative, we conduct a large-scale simulation study rigorously comparing single bootstrap, double bootstrap, and classical methods across multiple dimensions: sample size, confidence level, data-generating mechanisms, and statistical functionals. Results demonstrate that the double bootstrap consistently achieves superior coverage accuracy, stability, and robustness—particularly under small-sample and non-normal conditions—outperforming both classical approaches and the single bootstrap. We thus establish the double bootstrap as a principled, parsimonious, and high-performance paradigm for uncertainty quantification, providing both theoretical justification and empirical evidence to support its adoption in statistical education and applied practice.

Assessing bootstrap's potential to simplify statistical education and practiceComparing double bootstrap performance against traditional confidence interval techniquesEvaluating bootstrap as universal alternative for uncertainty quantification methods

Latest Papers

What's happening recently
View more

This study addresses the limitations of commonly reported point estimates—such as F1 scores—in text classification, which often lack reliable uncertainty quantification, particularly in settings involving small samples, rare classes, or nested data structures (e.g., texts nested within individuals). The authors systematically evaluate multiple confidence interval methods and propose a bootstrap-based F1 estimator augmented with pseudocount regularization. They further demonstrate that accurate inference in nested designs requires simultaneous adjustment of both effective sample size and degrees of freedom. Empirical results show that conventional Wald intervals suffer from undercoverage, whereas the recommended Agresti–Coull, Wilson, and hierarchical bootstrap methods substantially improve coverage accuracy, offering a more robust approach to uncertainty quantification in domains such as the social sciences.

classifier performanceconfidence intervalslarge language models

研究探讨了LLM-agent排行榜排名的实际意义,通过定义比较目标、检查共同支持并使用不确定性规则评估差异,揭示了排名相近时的不确定性及标签和规则对系统选择的影响。

estimand-awareLLM-agent leaderboardpairwise superiority

This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.

benchmarkingdataset selectionmodel ranking

This study addresses the potential pitfalls of directly applying item response theory (IRT)—originally designed for human assessment—to the evaluation of artificial intelligence systems, where mismatched data-generating mechanisms may compromise inference validity. It presents the first systematic evaluation of IRT’s applicability to large language model benchmarks, examining the feasibility, scalability, and reliability of four estimation approaches—marginal maximum likelihood, Markov chain Monte Carlo (MCMC), variational inference, and neural pseudo-twin estimators—across 18,000 simulated conditions. The findings reveal that classical methods are computationally infeasible at scale, while scalable alternatives introduce bias when the number of models is small or their ability distribution deviates from normality. The work quantifies, for the first time, the failure boundaries of IRT in AI evaluation and establishes required sample sizes and diagnostic criteria for its reliable application.

AI EvaluationBenchmarkingItem Response Theory

Hot Scholars

RT

Radu Timofte

Humboldt Professor for AI and Computer Vision, University of Würzburg
Computer VisionMachine LearningAICompression
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
LC

Linhan Cao

Shanghai Jiao Tong University
Image Quality Assessment Video Quality Assessment
ZW

Zongwei Wu

University of Würzburg | CNRS - Université de Bourgogne | ETH Zurich
Sensor FusionPerception