Score
Design and implement methods that compute uncertainty or confidence scores for model-produced samples at inference/test time and use those scores to select, discard, or reweight samples; this includes building scoring functions, selection rules, and pipelines for test-time filtering. Analyze and evaluate trade-offs between quality and coverage, and calibrate thresholds or weighting schemes to achieve desired output reliability and generation quality.
Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.
This work proposes an interpretable statistical inference framework that decomposes predictive scoring functions into three components—calibration error, discrimination ability, and uncertainty—applicable to multi-step-ahead point forecasts such as means and quantiles, and compatible with both smooth and nonsmooth scoring rules. Building on linear recalibration and integrating Mincer–Zarnowitz regression with asymptotic inference theory, the method delivers the first fully interpretable tripartite decomposition for general scoring functions, unifying and extending classical calibration tests and predictive performance evaluation. Empirical applications to inflation surveys and financial risk models reveal critical discrepancies obscured by aggregate scores, exposing a misalignment between backtesting practices and predictive accuracy in banking regulation, thereby substantially enhancing the informativeness and statistical power of forecast evaluation.
Traditional model evaluation relies on single-point metrics, failing to characterize performance stability and uncertainty. This paper proposes a small-sample (10–25 runs) uncertainty quantification framework tailored for high-reliability scenarios. It constructs empirical distributions of performance metrics via repeated stochastic experiments—encompassing random data splits, parameter initializations, and hyperparameter perturbations—and robustly estimates confidence intervals for metric quantiles using bias-corrected nonparametric bootstrap combined with quantile regression. To our knowledge, this is the first systematic approach enabling reliable confidence interval estimation for diverse metrics—including accuracy, F1-score, and MAE—in both classification and regression tasks under small-sample regimes. The method achieves high coverage (>90%) while maintaining narrow interval widths, thereby significantly improving robustness in model selection and enhancing decision-making credibility across multiple benchmark datasets.
Current sample size calculations for clinical prediction models ensure only that performance metrics meet target values *on average*, neglecting sampling variability—resulting in unstable model performance and low probability of achieving acceptable performance (PrAP) in practice. This paper proposes a novel sample size determination framework centered on PrAP, formally adopting “probability of attaining acceptable performance” as the primary statistical objective—replacing conventional expectation-based approaches. Through simulation studies and analytical derivations, we develop robust methods for estimating calibration slope in binary outcome settings, implemented in the R package `samplesizedev`. Results demonstrate that conventional methods yield PrAPs typically below 60%, whereas our approach consistently achieves PrAPs exceeding 80%, with particularly pronounced gains when fewer predictors are included. This substantially improves model reliability and reproducibility.
Deep neural network training suffers from high sensitivity to random seeds due to stochastic optimization, hindering reliable assessment of true generalization performance. To address this, we propose a robust nonparametric hypothesis testing framework. Its core innovation is a novel model similarity metric—the α-truncation level—which quantifies training variability and determines the minimum number of independent training runs required for stable ensembling. Unlike conventional metrics such as accuracy or expected calibration error (ECE), the α-truncation level does not rely on modeling the null distribution and is inherently sensitive to training instability. Experiments demonstrate that it detects training uncertainty earlier and more consistently than validation accuracy, churn, and ECE. Moreover, in transfer learning settings, it effectively guides random seed selection, significantly improving the reliability of performance evaluation.
Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.
本文针对模型信息不足时的可靠性集合估计问题,提出一种结合校准过程和自适应设计的统一框架,以降低决策风险并控制误包含风险。
本文提出一种基于置信区域的筛选框架,用于解决模拟系统可接受性问题,保证高概率筛选出所有或每个可接受系统,并支持并行化。
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
论文提出Evidence-Calibration-Stability框架,通过区分证据、校准和稳定性来解决模型不确定性下的假设检验问题。