Score
Designs and implements resampling procedures and evaluation pipelines that preserve, correct, or measure a model's probability calibration when calibration data are subsampled or class frequencies are altered. Builds analyses and diagnostics that quantify how different sampling schemes affect calibration estimates, align empirical sampling frequencies with predicted probabilities, and mitigate sampling mismatch for underrepresented outcomes.
This paper addresses the reliability of calibration evaluation for machine learning models, identifying systematic biases in the widely used Expected Calibration Error (ECE) under distributional shift and varying binning strategies. Methodologically, it clarifies the logical hierarchy among multi-level calibration definitions, and systematically exposes ECE’s limitations through visualization, binning-based statistical analysis, and theoretical derivation—demonstrating its failure to satisfy key requirements of robustness and consistency in calibration assessment. Building on this critique, the paper introduces and explicates emerging calibration paradigms—including distribution-level and instance-level calibration—alongside their corresponding evaluation methodologies, thereby constructing a rigorous, interpretable, and practice-oriented calibration knowledge framework. The results equip researchers with principled guidance for selecting appropriate evaluation metrics and advance calibration assessment from ad hoc, heuristic practices toward standardization and formalization.
Current sample size calculations for clinical prediction models ensure only that performance metrics meet target values *on average*, neglecting sampling variability—resulting in unstable model performance and low probability of achieving acceptable performance (PrAP) in practice. This paper proposes a novel sample size determination framework centered on PrAP, formally adopting “probability of attaining acceptable performance” as the primary statistical objective—replacing conventional expectation-based approaches. Through simulation studies and analytical derivations, we develop robust methods for estimating calibration slope in binary outcome settings, implemented in the R package `samplesizedev`. Results demonstrate that conventional methods yield PrAPs typically below 60%, whereas our approach consistently achieves PrAPs exceeding 80%, with particularly pronounced gains when fewer predictors are included. This substantially improves model reliability and reproducibility.
This work addresses the challenge that probabilistic programs generated by language models often suffer from statistical misspecifications—such as incorrect likelihoods, priors, or parameterizations—that are difficult to detect with conventional unit tests. The paper introduces, for the first time, Bayesian calibration as a central criterion for assessing the correctness of probabilistic programs and proposes a fully unsupervised, reference-free framework for their detection and repair. By integrating Bayesian validation techniques—including posterior predictive checks, simulation-based calibration (SBC), sampling diagnostics (e.g., $\hat{R}$, divergences, effective sample size), and held-out predictive log density—the method generates feedback signals to drive an iterative repair loop within large language models. Evaluated on 200 instances, the approach achieves detection AUCs of 0.97 with reference programs and 62–78% without, substantially outperforming unit testing; repair success rates reach 92% and 100% using GPT-5.1 and Claude, respectively.
This study addresses the adverse impact of resampling methods—such as SMOTE and random undersampling—on the calibration of tree-based ensemble models under class imbalance. While these techniques improve classification performance, they degrade probability calibration, thereby compromising the reliability of decisions that depend on predicted probabilities. The work systematically evaluates this effect and quantifies, for the first time, that SMOTE increases the expected calibration error (ECE) by an average of 0.009, whereas random undersampling under high imbalance elevates ECE to as much as 0.395. It further demonstrates that standard prior-probability correction is ineffective for SMOTE, necessitating data-driven post-hoc calibration. Experiments show that applying Platt or isotonic regression reduces ECE by up to 66% with negligible AUC degradation (only 0.002), underscoring the necessity and efficacy of post-calibration in imbalanced learning scenarios.
This paper addresses the problem that conventional calibration evaluation of deep learning models is vulnerable to spurious recalibration—i.e., trivial post-hoc adjustments that improve calibration metrics without enhancing generalization. To tackle this, we propose a novel joint evaluation paradigm integrating calibration and generalization. First, we derive a Bregman-divergence-based decomposition of calibration error, establishing the first theoretical connection between calibration metrics and generalization objectives (e.g., negative log-likelihood). Second, we design a new reliability diagram that jointly visualizes calibration bias and estimated generalization error. Third, we characterize multiple “pseudo-optimal” calibration phenomena and provide theoretically grounded, detectable criteria for identifying trivial recalibration. Experiments on standard benchmarks demonstrate that our approach significantly improves model diagnostic capability, yielding a more reliable and interpretable evaluation framework for calibration research.
This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.
该研究提出CORD方法,通过修复校准后的概率向量来保持原始预测不变,同时改善了模型的校准性能。
本文提出了一种新的度量方法rankECE,通过比较预测概率相近的点来更准确地估计模型的校准误差,以解决现有ECE估计方法不准确的问题。
This work introduces calibration as a novel dimension for evaluating fine-grained subtype robustness, examining model reliability when encountering subtypes unseen during training yet belonging to known coarse-grained categories. Systematic evaluation across ImageNet, BREEDS, iNaturalist, and CIFAR-100 on five mainstream architectures reveals that models exhibit severe overconfidence: their confidence fails to decrease appropriately when accuracy drops due to exposure to novel subtypes. This miscalibration under semantic shift is markedly more pronounced than under common image corruptions. Existing recalibration techniques and out-of-distribution detection methods only partially mitigate the issue, underscoring the necessity of treating calibration as an independent and essential criterion in robustness assessment.
论文提出Evidence-Calibration-Stability框架,通过区分证据、校准和稳定性来解决模型不确定性下的假设检验问题。