Score
Designs, implements, and evaluates algorithms and metrics that assign numerical or categorical confidence estimates to model outputs and dataset labels (including calibration and scoring methods). Builds and analyzes selection and weighting procedures that use those confidence estimates to choose, weight, or filter samples and to assess/report reliability and its impact on downstream performance and trustworthiness.
Probabilistic outputs of AI models often exhibit miscalibration—i.e., predicted confidence scores poorly reflect true accuracy—hindering their reliable deployment in safety-critical applications and ensemble systems. Method: This paper presents a systematic survey of probabilistic calibration evaluation methods for classification and object detection models. Grounded in statistical assessment theory, it unifies diverse approaches—including reliability diagrams, Brier score, expected calibration error (ECE), maximum calibration error (MCE), Kolmogorov–Smirnov test, ROC-based metrics, and IoU-aware measures—within a coherent framework covering binary, multiclass, and detection tasks. Contribution: We propose the first taxonomy of calibration metrics, categorizing 82 existing measures into four families: point-wise, binning-based, kernel/curve-based, and cumulative. Additionally, we introduce the first structured calibration metric knowledge base, enabling rapid metric selection, implementation, and comparative analysis—thereby establishing new interpretable and quantifiable benchmarks for trustworthy AI.
Existing confidence calibration methods predominantly rely on statistical fitting, neglecting the underlying prior distribution governing calibration curves. This work proposes a novel calibration framework based on the Binomial Process (BPM), the first to model calibration data as a binomial process. We theoretically establish its Lipschitz continuity and high sample efficiency—requiring only $3/B$ samples compared to histogram-based methods (where $B$ is the number of bins). Our approach jointly incorporates prior knowledge and empirical observations, fitting a continuous calibration curve via maximum likelihood estimation and joint optimization. We further introduce the Total Calibration Error (TCE$_{ ext{pm}}$), a consistent and unbiased metric for calibration error assessment. Extensive experiments on both synthetic and real-world datasets demonstrate that our method significantly outperforms state-of-the-art approaches in calibration accuracy, robustness under limited samples, and consistency of error estimation.
Traditional model evaluation relies on single-point metrics, failing to characterize performance stability and uncertainty. This paper proposes a small-sample (10–25 runs) uncertainty quantification framework tailored for high-reliability scenarios. It constructs empirical distributions of performance metrics via repeated stochastic experiments—encompassing random data splits, parameter initializations, and hyperparameter perturbations—and robustly estimates confidence intervals for metric quantiles using bias-corrected nonparametric bootstrap combined with quantile regression. To our knowledge, this is the first systematic approach enabling reliable confidence interval estimation for diverse metrics—including accuracy, F1-score, and MAE—in both classification and regression tasks under small-sample regimes. The method achieves high coverage (>90%) while maintaining narrow interval widths, thereby significantly improving robustness in model selection and enhancing decision-making credibility across multiple benchmark datasets.
Traditional model confidence sets rely on the fixed-sample assumption, rendering them inadequate for continuous model selection and uncertainty quantification under dynamic data streams. To address this, we introduce the first sequential extension of model confidence sets, proposing a dynamic model screening framework grounded in e-processes, confidence sequences, and sequential hypothesis testing. Our method requires no prespecified stopping rule and delivers, at any time, a nonasymptotic, time-uniform confidence set that provably covers the true optimal model subset with guaranteed nominal coverage probability. Unlike static approaches, it substantially enhances statistical robustness and real-time adaptability in online model evaluation. The framework provides both theoretical guarantees and practical tools for trustworthy model selection in streaming data analysis.
This work addresses the challenge of quantifying prediction uncertainty in generative biomolecular design, where feedback covariate shift undermines conventional uncertainty estimation. We propose the first conformal prediction framework tailored to closed-loop design paradigms. Departing from standard i.i.d. assumptions, our method imposes no structural constraints on either the design algorithm or the regression model, delivering finite-sample statistically valid confidence sets for arbitrary black-box design pipelines. Key innovations include quantile-regression-driven adaptive conformal prediction, explicit modeling of feedback-induced distributional shift, and robust error calibration. Evaluated on protein and small-molecule design tasks, our approach achieves ≥94.8% empirical coverage at the 95% nominal confidence level—substantially outperforming standard conformal methods (which drop to as low as 72%)—while maintaining high predictive accuracy.
In causal subgroup identification, conventional methods suffer from high estimation noise in conditional average treatment effect (CATE) estimation and multiplicity issues arising from two-stage procedures. To address these challenges, this paper proposes the Global Adaptive Treatment Effect Sets (GATES) uniform confidence band method. Grounded in randomized trial design and empirical process theory, GATES provides finite-sample, model-agnostic global statistical guarantees for CATE estimates produced by arbitrary black-box machine learning models—without requiring parametric assumptions or resampling. It enables rigorous, threshold-agnostic identification of credible subgroups exhibiting clinically meaningful treatment effects. Empirically, GATES maintains nominal coverage even in small samples (n = 100), substantially improving the reliability of subgroup inference. Applied to a late-stage prostate cancer clinical trial, it robustly identifies a clinically significant “exceptional responder” subgroup. This work establishes a verifiable, statistically principled paradigm for causal subgroup discovery in precision medicine.
This work proposes a novel representation learning framework that addresses the limited representational capacity of existing methods in complex scenes by integrating adaptive multi-scale fusion with contrastive learning. The approach dynamically aggregates multi-level features and incorporates a structure-aware contrastive loss, thereby enhancing the model’s ability to jointly capture fine-grained semantics and global contextual information. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art methods across multiple benchmark datasets, achieving substantial improvements in both accuracy and robustness. These results establish a promising new direction for unsupervised and semi-supervised representation learning.
This study addresses the urgent need for high-accuracy demographic and socioeconomic indicators under budget constraints by developing an AI-driven framework that ensures credibility, privacy preservation, and statistical validity. Integrating the United Nations’ official statistics principles with the HLG-MOS algorithmic quality standards, the work proposes a verifiable and auditable AI assessment checklist and guides the development of the MiniMax hierarchical Bayesian sampling algorithm. Validated through Monte Carlo simulations, synthetic populations, and real census microdata, the approach achieves an 80% reduction in sample size while meeting precision requirements on synthetic labor data and enables a 90% sample reduction on the 2021 Australian Census data, yielding national point estimates with errors below 1%. This significantly enhances the efficiency and reliability of official statistics production.
This work addresses the limitations of the standard Expected Calibration Error (ECE), which struggles to effectively capture overconfidence risks at high confidence levels and fails to evaluate the discriminative power of confidence scores with respect to prediction correctness. To overcome these issues, the authors propose the Calibrated Size Ratio (CSR) as a more sensitive calibration metric and introduce the risk probability \(P_{\text{risk}}\) to quantify overconfidence. Furthermore, they systematically extend confidence-weighting mechanisms to various classification metrics for the first time, yielding novel measures such as cwA and cwAUC to assess the discriminative ability of confidence estimates. Theoretical analysis and extensive experiments across 15 real-world and synthetic datasets demonstrate that CSR consistently exhibits superior sensitivity and specificity across diverse calibration scenarios, validating the effectiveness and robustness of the proposed approach.
In A/B testing, control variates and regression adjustment are widely used variance reduction techniques, yet their theoretical relationship remains unclear, their methodological frameworks are disjointed, and both have long been confined to design-driven paradigms. Method: This paper establishes, for the first time, a formal equivalence between these two approaches and proposes a novel grouped coefficient estimation method that unifies design-based and model-based estimation frameworks—enabling a paradigm shift from design-driven to model-driven inference. Contribution/Results: Theoretical analysis demonstrates improved estimation accuracy and statistical power. Empirical validation on millions of real-world experiments at ByteDance confirms efficacy: the proposed method has been fully deployed in its online experimentation platform, yielding an average 12.3% increase in statistical significance and a 19.6% improvement in detection sensitivity.
Current sample size calculations for clinical prediction models ensure only that performance metrics meet target values *on average*, neglecting sampling variability—resulting in unstable model performance and low probability of achieving acceptable performance (PrAP) in practice. This paper proposes a novel sample size determination framework centered on PrAP, formally adopting “probability of attaining acceptable performance” as the primary statistical objective—replacing conventional expectation-based approaches. Through simulation studies and analytical derivations, we develop robust methods for estimating calibration slope in binary outcome settings, implemented in the R package `samplesizedev`. Results demonstrate that conventional methods yield PrAPs typically below 60%, whereas our approach consistently achieves PrAPs exceeding 80%, with particularly pronounced gains when fewer predictors are included. This substantially improves model reliability and reproducibility.