Score
Designs and implements an integrated system that produces predictive probabilities and then evaluates and adjusts their calibration both at the model and instance (behavioral) level, including post-hoc recalibration, calibration-aware scoring, and monitoring. Builds pipelines that link predictive outputs to calibration analysis and feature-level explanations, and computes model- and example-level calibration metrics to detect, quantify, and guide remediation of miscalibration.
Probabilistic outputs of AI models often exhibit miscalibration—i.e., predicted confidence scores poorly reflect true accuracy—hindering their reliable deployment in safety-critical applications and ensemble systems. Method: This paper presents a systematic survey of probabilistic calibration evaluation methods for classification and object detection models. Grounded in statistical assessment theory, it unifies diverse approaches—including reliability diagrams, Brier score, expected calibration error (ECE), maximum calibration error (MCE), Kolmogorov–Smirnov test, ROC-based metrics, and IoU-aware measures—within a coherent framework covering binary, multiclass, and detection tasks. Contribution: We propose the first taxonomy of calibration metrics, categorizing 82 existing measures into four families: point-wise, binning-based, kernel/curve-based, and cumulative. Additionally, we introduce the first structured calibration metric knowledge base, enabling rapid metric selection, implementation, and comparative analysis—thereby establishing new interpretable and quantifiable benchmarks for trustworthy AI.
Existing studies lack rigorous theoretical characterization of posterior calibration methods—such as Platt scaling and isotonic regression—particularly regarding their dependence on feature quality, generalizability across models and datasets, and convergence behavior and robustness under finite-sample regimes. Method: We establish a unified theoretical framework for these two dominant calibration paradigms, deriving the first non-asymptotic guarantees on convergence rates, computational complexity, and explicit sample-size dependencies. Our analysis quantifies the relationship between feature informativeness and calibration robustness. Results: Through synthetic experiments and extensive empirical evaluation across diverse model architectures and benchmark datasets, we validate our theoretical findings. The results yield actionable guidance: isotonic regression is preferable under low signal-to-noise ratios or limited samples, whereas Platt scaling exhibits superior robustness in high-dimensional sparse feature settings. Our work provides interpretable, reusable principles for uncertainty calibration in practical machine learning systems.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
To address the degradation of prediction probability calibration in deployed image classification models due to concept drift, this paper proposes an online calibration monitoring method that requires no access to model internals—only predicted probabilities and ground-truth labels. Our approach introduces, for the first time, a Cumulative Sum (CUSUM) control chart with dynamic control limits into calibration monitoring. It computes cumulative deviations of calibration error over time and adaptively adjusts detection thresholds to enable early warning of calibration loss. Compared to static-threshold methods, our framework significantly enhances sensitivity to temporal distribution shifts and accelerates response to emerging miscalibration. We validate its effectiveness and robustness across multiple image classification benchmarks under diverse concept drift scenarios. The proposed method establishes a scalable, black-box-compatible paradigm for trustworthy model deployment, enabling continuous, lightweight calibration assessment without architectural or training modifications.
In safety-critical applications, evaluating uncertainty calibration of regression models is hindered by inconsistent metric definitions, conflicting assumptions, and incomparable scales—impeding interpretability and reproducibility. This work systematically surveys and categorizes existing calibration metrics, then conducts a model-agnostic benchmark across real-world, synthetic, and manually miscalibrated datasets. We empirically demonstrate—for the first time—that most metrics yield contradictory or even opposing conclusions for identical calibration states, confirming that metric choice critically influences research outcomes. To address this, we propose ENCE (Expected Normalized Calibration Error) and CWC (Weighted Coverage Confidence) as more robust and stable primary metrics. Experiments across diverse scenarios show that ENCE and CWC exhibit superior consistency, strong resilience to noise and distribution shifts, and enhanced interpretability. Our findings establish a reproducible methodological foundation for uncertainty calibration evaluation in regression.
This paper addresses the problem that conventional calibration evaluation of deep learning models is vulnerable to spurious recalibration—i.e., trivial post-hoc adjustments that improve calibration metrics without enhancing generalization. To tackle this, we propose a novel joint evaluation paradigm integrating calibration and generalization. First, we derive a Bregman-divergence-based decomposition of calibration error, establishing the first theoretical connection between calibration metrics and generalization objectives (e.g., negative log-likelihood). Second, we design a new reliability diagram that jointly visualizes calibration bias and estimated generalization error. Third, we characterize multiple “pseudo-optimal” calibration phenomena and provide theoretically grounded, detectable criteria for identifying trivial recalibration. Experiments on standard benchmarks demonstrate that our approach significantly improves model diagnostic capability, yielding a more reliable and interpretable evaluation framework for calibration research.
This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.
This study addresses the capability degradation and goal-contract invalidation of AI agents caused by shifting conditions during model transfer, cross-domain deployment, or scaling. To mitigate these issues, we propose a three-layer interactive calibration methodology that decouples trainable policies from frozen configurations, establishing a multidimensional constraint system encompassing information preservation, execution framework adaptation, and user acceptance. Technically, the approach integrates semantic checkpoint repair, tool substitution, local replanning, and output contract enforcement mechanisms. For evaluation, we introduce a factorial testing holdout system to prevent aggregated gains from masking localized failures. This work provides a methodological framework enabling the coexistence of shared standards and market-specific adapters in global e-commerce scenarios, with empirical validation reserved for future research.
本文提出了一种新的度量方法rankECE,通过比较预测概率相近的点来更准确地估计模型的校准误差,以解决现有ECE估计方法不准确的问题。
Traditional Bayesian calibration struggles in dynamic systems to disentangle model parameters from discrepancy terms and is ill-equipped to handle both gradual drifts and abrupt shifts, often being confined to offline settings. This work proposes the Bayesian Recursive Projection Calibration (BRPC) framework, which extends projection-based calibration to online scenarios for the first time. BRPC ensures identifiability and tracks gradual changes by decoupling parameter updates from discrepancy modeling via Gaussian processes, while incorporating a theoretically grounded restart mechanism coupled with online change detection to respond to sudden shifts. Experimental results on synthetic data and industrial process simulations demonstrate that BRPC significantly outperforms sliding-window Bayesian calibration and data assimilation baselines, achieving higher accuracy under gradual drift and maintaining robustness during abrupt changes.
This study addresses the longstanding disconnect between student academic performance prediction and metacognitive calibration by proposing a Unified Behavioral Prediction and Calibration Analysis Pipeline (UBP-CAP). Integrating prediction, calibration assessment, and variance decomposition modules, UBP-CAP leverages multimodal telemetry data to simultaneously predict response accuracy and quantify metacognitive bias. The work introduces the Prediction-Explanation Discrepancy Index (PEDI) to measure feature consistency between predictive and explanatory models and employs cross-validated generalized linear mixed-effects models (GLMMs) to uncover the context-dependence of calibration bias. Empirical results show that logistic regression (AUC = 0.903) outperforms LightGBM; students exhibit significantly higher calibration error (ECE = 0.109) than the model (ECE = 0.068); GLMM analysis yields an intraclass correlation coefficient (ICC) of 0.123, indicating calibration is predominantly context-driven; and PEDIcos = 0.081 reveals high alignment between prediction and explanation features.