compare models statistically

Designs and executes controlled comparative evaluations that quantify differences between models by selecting and computing appropriate comparison metrics and experimental protocols (including cross-model and human–model comparisons). Builds and runs the statistical analyses and uncertainty quantification needed to support claims — e.g., paired and permutation tests, Bayesian model comparison, confidence intervals, clustered standard errors, and power/sample-size calculations — and reports publication-ready comparative conclusions with correct significance assessment.

comparemodelsstatistically

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$216K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Bayesian sample size calculations for external validation studies of risk prediction models

Apr 22, 2025
MS
Mohsen Sadatsafavi
🏛️ the University of British Columbia | British Columbia Centre for Disease Control | Maastricht University | KU Leuven | University of Birmingham | National Institute for Health and Care Research

Conventional sample size calculations for external validation of risk prediction models rely on fixed prior assumptions about model performance (e.g., calibration, discrimination) and net benefit (NB), failing to capture real-world uncertainty; moreover, traditional precision-oriented approaches—based on confidence interval width—bear weak relevance to clinical utility (NB). Method: This paper introduces, for the first time, a systematic Bayesian framework for external validation sample size determination. It constructs a joint risk–outcome distribution from prior performance summaries and proposes a multi-objective Bayesian sample size rule that jointly optimizes expected precision, assurance probability, optimal NB identification, and expected value of sample information (EVSI). Contribution/Results: The method enables decision-making under quantified uncertainty, enhancing statistical robustness and clinical relevance. Applied to external validation of a COVID-19 deterioration risk model, it improves resource efficiency and reliability of validation conclusions.

Addressing limitations of conventional precision-based inference for net benefitBayesian sample size calculations for uncertain model performanceProposing rules for sample size based on assurance probabilities and EVSI

This study addresses the critical yet underexamined practice of “borrowing treatment effects” (rather than individual patient data) in Bayesian treatment effect extrapolation. We conduct the first large-scale frequentist simulation study to systematically evaluate the operating characteristics of four methods: conditional power priors, robust mixture priors, post-hoc pooling, and p-value power priors. Results demonstrate that conditional power priors and robust mixture priors achieve the best overall performance in terms of success probability, bias control, and nominal coverage of credible intervals. In contrast, post-hoc pooling and p-value power priors exhibit substantial upward bias and severe undercoverage. This work fills a key methodological gap by providing the first frequentist evaluation framework for Bayesian extrapolation at the treatment-effect level. It delivers empirical evidence and practical guidance for selecting appropriate extrapolation strategies in confirmatory trials.

Assessing operating characteristics in confirmatory trial contextsComparing performance of different Bayesian borrowing techniquesEvaluating Bayesian methods for treatment effect extrapolation

Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.

Bayesian decision proceduresexperimental designgeneralized posteriors

In A/B testing, control variates and regression adjustment are widely used variance reduction techniques, yet their theoretical relationship remains unclear, their methodological frameworks are disjointed, and both have long been confined to design-driven paradigms. Method: This paper establishes, for the first time, a formal equivalence between these two approaches and proposes a novel grouped coefficient estimation method that unifies design-based and model-based estimation frameworks—enabling a paradigm shift from design-driven to model-driven inference. Contribution/Results: Theoretical analysis demonstrates improved estimation accuracy and statistical power. Empirical validation on millions of real-world experiments at ByteDance confirms efficacy: the proposed method has been fully deployed in its online experimentation platform, yielding an average 12.3% increase in statistical significance and a 19.6% improvement in detection sensitivity.

Analyzing statistical properties and theoretical connections between frameworksBridging control variates and regression adjustment methodsProviding guidance for variance reduction in A/B testing

On the handling of method failure in comparison studies

Aug 21, 2024
MW
Milena Wunsch
🏛️ LMU Munich | Munich Center for Machine Learning | Department of Statistics | MRC Clinical Trials Unit | UCL

In methodological comparative studies, algorithmic failures—such as non-convergence or absence of output—preclude performance evaluation, yet existing literature lacks standardized guidelines for handling such failures, often overlooking or misapplying failure mitigation strategies. Method: We systematically analyze failure causes and risks of improper handling, critically examine prevalent censoring and imputation strategies for their statistical biases, and propose the principle of “context-adapted failure fallback,” establishing a framework grounded in empirically feasible fallback mechanisms. Through statistical modeling, failure root-cause diagnosis, and cross-domain empirical analysis, we identify widespread deficiencies in published studies’ failure handling practices. Contribution/Results: Two representative case studies demonstrate that inappropriate failure handling significantly distorts method rankings and undermines conclusion validity. Our work bridges critical theoretical and practical gaps in the principled treatment of algorithmic failures in empirical methodology research.

Addressing method failure handling in comparison studiesProviding guidance on proper failure interpretation and reportingRecommending realistic fallback strategies for method failures

Latest Papers

What's happening recently
View more

Estimating the functional relationship between a continuous exposure and a binary outcome is challenging when covariates are measured with error. This study presents the first systematic evaluation of Simulation-Extrapolation, Regression Calibration, multiple imputation, and Bayesian correction methods, each coupled with flexible modeling techniques—including B-splines, P-splines, and fractional polynomials—within a multi-team, fully blinded, neutral simulation framework. By generating 155 distinct simulation scenarios and repeated samples, the research quantifies the bias and variance of each approach, revealing their relative strengths and limitations. The findings not only inform method selection under measurement error but also demonstrate the feasibility and value of this neutral comparative paradigm for rigorous methodological assessment.

covariate adjustmentexposure-outcome relationshipfunctional form

Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.

computational costfixed-size benchmarksmodel evaluation

Current performance evaluation metrics—such as accuracy and F1 score—are typically reported as point estimates, ignoring the uncertainty induced by data clustering structures. This oversight often leads to underestimation of variability and potentially misleading model comparisons. To address this, this work proposes a unified framework that expresses a broad class of performance metrics as smooth functionals of the confusion matrix probabilities. By integrating a cluster-robust sandwich variance estimator, the framework enables valid confidence interval construction, hypothesis testing, and paired model comparison. It represents the first systematic application of cluster-robust inference to predictive performance evaluation, accommodating both binary and multiclass settings, and further provides asymptotic theory–based methods for power and sample size calculations. Simulations demonstrate that the proposed approach achieves near-nominal coverage across diverse dependence structures and substantially outperforms conventional methods that ignore clustering; real-data analyses confirm that accounting for clustering can materially alter evaluation conclusions.

clustered datadependent datamodel evaluation

Hot Scholars

HY

Hitomi Yanaka

The University of Tokyo, RIKEN
Natural Language ProcessingSemantics
AF

Alessio Ferrari

Lecturer, UCD; Senior Research Scientist, ISTI CNR
Natural Language ProcessingRequirements EngineeringRequirements ElicitationFormal Methods
SM

Sébastien Marcel

Senior researcher ( Idiap research institute ) and Professor ( University of Lausanne )
AIbiometricssecurity and privacyanti-spoofing and deepfakes
PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy
YL

Yankai Lin

Associate Professor (Tenure Track), Gaoling School of AI, Renmin University of China
Natural Language ProcessingLarge Language Models