Score
Designs and implements methods and tools to manage leaderboards and perform statistically rigorous evaluation of models and benchmarks. This includes computing per-task bootstrap confidence intervals, running paired-bootstrap significance tests, applying item-response-theory discrimination analyses, and measuring cross-leaderboard correlations to enable robust model comparison and uncertainty-aware ranking.
This work addresses a critical limitation in current model leaderboard evaluation methods, which often overlook task-level performance uncertainty and variability, thereby yielding unreliable global rankings. To remedy this, the authors propose a hierarchical framework that enables principled aggregation of uncertainty from the task level to the leaderboard level. By integrating task-wise pairwise comparisons with conformal prediction, the method produces statistically valid confidence intervals for both individual task rankings and overall model rankings. Empirical validation on synthetic data as well as real-world benchmarks—TabArena and PromptEval (MMLU)—demonstrates the approach’s statistical validity and informativeness. The resulting rankings are not only reliable but also explicitly account for uncertainty, offering a robust foundation for evaluating and comparing models on new tasks.
This study addresses the problem of uncontrolled confidence interval error rates in dynamic evaluation of model leaderboards caused by repeated testing or early stopping. We propose an anytime-valid confidence sequence method for full-model ranking that combines betting e-processes for pairwise comparisons with closed testing to integrate all possible orderings. This approach constructs, for the first time under score-dependent conditions, confidence sequences that permit arbitrary stopping rules without inflating error rates. Simulations and experiments on public datasets demonstrate that the proposed method supports real-time monitoring and adaptive stopping, significantly reducing computational costs while sacrificing only minimal statistical power compared to fixed-sample approaches.
This study addresses the stability assessment of group-specific ranking patterns and nonparametric inference on population-mean ordinal relationships in multivariate survey/scoring data. We propose the first hierarchical bootstrap framework for ordinal hypothesis testing, which approximates the null distribution without distributional assumptions. We introduce the *non-containment index*—a robust, interpretable measure quantifying ranking stability across groups—and leverage it for outlier response detection and significance testing of inter-group ranking differences. The method integrates hierarchical resampling, nonparametric stability analysis, and resampling-based ordinal inference, unifying descriptive and inferential capabilities. Evaluated in AI fairness auditing and questionnaire analysis, it demonstrates high sensitivity and reliability. Our approach establishes a novel, assumption-free, robust, and interpretable statistical paradigm for ordinal data analysis.
This paper addresses the challenge of quantifying uncertainty in cross-validation (CV) performance estimates—particularly the difficulty of distinguishing true performance differences from random fluctuations during model comparison. We propose an efficient and robust bootstrap-based method that decomposes the variance of CV estimates using a random-effects model, enabling valid statistical inference on CV performance differences without strong modeling assumptions. Compared to standard bootstrap, our approach substantially reduces computational cost while yielding confidence intervals with accurate coverage probability and high statistical power. Extensive evaluations—including simulations and real-world applications across classification, regression, and time-series tasks—demonstrate its strong robustness and generalizability under small-sample settings, non-independent CV folds, and heterogeneous data distributions. The method provides reliable uncertainty quantification to support principled model selection.
This paper addresses the complexity and high pedagogical/practical barriers associated with conventional uncertainty quantification methods—such as standard errors, confidence intervals, and hypothesis tests—in statistical inference. To evaluate the potential of nonparametric bootstrap as a unified alternative, we conduct a large-scale simulation study rigorously comparing single bootstrap, double bootstrap, and classical methods across multiple dimensions: sample size, confidence level, data-generating mechanisms, and statistical functionals. Results demonstrate that the double bootstrap consistently achieves superior coverage accuracy, stability, and robustness—particularly under small-sample and non-normal conditions—outperforming both classical approaches and the single bootstrap. We thus establish the double bootstrap as a principled, parsimonious, and high-performance paradigm for uncertainty quantification, providing both theoretical justification and empirical evidence to support its adoption in statistical education and applied practice.
This study addresses the limitations of commonly reported point estimates—such as F1 scores—in text classification, which often lack reliable uncertainty quantification, particularly in settings involving small samples, rare classes, or nested data structures (e.g., texts nested within individuals). The authors systematically evaluate multiple confidence interval methods and propose a bootstrap-based F1 estimator augmented with pseudocount regularization. They further demonstrate that accurate inference in nested designs requires simultaneous adjustment of both effective sample size and degrees of freedom. Empirical results show that conventional Wald intervals suffer from undercoverage, whereas the recommended Agresti–Coull, Wilson, and hierarchical bootstrap methods substantially improve coverage accuracy, offering a more robust approach to uncertainty quantification in domains such as the social sciences.
研究探讨了LLM-agent排行榜排名的实际意义,通过定义比较目标、检查共同支持并使用不确定性规则评估差异,揭示了排名相近时的不确定性及标签和规则对系统选择的影响。
This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.
This study addresses the potential pitfalls of directly applying item response theory (IRT)—originally designed for human assessment—to the evaluation of artificial intelligence systems, where mismatched data-generating mechanisms may compromise inference validity. It presents the first systematic evaluation of IRT’s applicability to large language model benchmarks, examining the feasibility, scalability, and reliability of four estimation approaches—marginal maximum likelihood, Markov chain Monte Carlo (MCMC), variational inference, and neural pseudo-twin estimators—across 18,000 simulated conditions. The findings reveal that classical methods are computationally infeasible at scale, while scalable alternatives introduce bias when the number of models is small or their ability distribution deviates from normality. The work quantifies, for the first time, the failure boundaries of IRT in AI evaluation and establishes required sample sizes and diagnostic criteria for its reliable application.
研究通过审计254个SWE-bench提交,发现编码代理排行榜上的微小差异不能准确反映系统优劣,建议报告比较集分辨率和模型-框架来源。