Score
Designs, implements, or analyzes methods that rescale model logits or softmax inputs using a global or learned temperature parameter to calibrate output probabilities, align probability distributions across models or runs, and control the effective logit scale; evaluates the impact of such rescaling on calibration metrics, cross-model probability alignment, and training or inference dynamics.
Neural networks commonly exhibit systematic overconfidence, posing safety risks in critical deployment scenarios. Existing calibration methods face a bias-variance trade-off: global temperature scaling is efficient but suffers from high bias; more expressive methods incur high variance due to noise in high-dimensional logits and scarcity of calibration data. This paper proposes LogitGap-TS—a lightweight, data-efficient, sample-aware calibration method that uses the gap between the top-two logits as a denoised scalar signal to dynamically estimate per-sample temperature parameters. We further introduce the SoftECE loss, enabling robust bias-variance optimization under limited calibration data. Evaluated across multiple models and datasets, LogitGap-TS achieves state-of-the-art calibration performance with significantly fewer parameters than existing approaches. It converges stably using only a minimal number of calibration samples—e.g., as few as 32—demonstrating exceptional data efficiency and practicality for safety-critical applications.
Existing post-hoc calibration methods for multi-class classifiers—particularly logistic regression–based approaches—suffer from overfitting due to excessive parameters and limited calibration data. Method: We propose Structured Matrix Scaling (SMS), a principled calibration framework built upon multinomial logistic regression, incorporating structured matrix regularization (e.g., low-rank or diagonal-plus-low-rank constraints), cross-class parameter sharing, and robust feature preprocessing. This design simultaneously enhances expressivity and controls variance. Contribution/Results: SMS theoretically overcomes the representational limitations of temperature scaling and standard matrix scaling, achieving a superior bias–variance trade-off. Extensive experiments across diverse models and datasets demonstrate that SMS significantly outperforms existing logistic regression–based calibration methods, while exhibiting strong scalability and practicality. The implementation is open-sourced, establishing a new efficient and robust benchmark for probabilistic calibration.
This study addresses the capability degradation and goal-contract invalidation of AI agents caused by shifting conditions during model transfer, cross-domain deployment, or scaling. To mitigate these issues, we propose a three-layer interactive calibration methodology that decouples trainable policies from frozen configurations, establishing a multidimensional constraint system encompassing information preservation, execution framework adaptation, and user acceptance. Technically, the approach integrates semantic checkpoint repair, tool substitution, local replanning, and output contract enforcement mechanisms. For evaluation, we introduce a factorial testing holdout system to prevent aggregated gains from masking localized failures. This work provides a methodological framework enabling the coexistence of shared standards and market-specific adapters in global e-commerce scenarios, with empirical validation reserved for future research.
This paper addresses the problem that conventional calibration evaluation of deep learning models is vulnerable to spurious recalibration—i.e., trivial post-hoc adjustments that improve calibration metrics without enhancing generalization. To tackle this, we propose a novel joint evaluation paradigm integrating calibration and generalization. First, we derive a Bregman-divergence-based decomposition of calibration error, establishing the first theoretical connection between calibration metrics and generalization objectives (e.g., negative log-likelihood). Second, we design a new reliability diagram that jointly visualizes calibration bias and estimated generalization error. Third, we characterize multiple “pseudo-optimal” calibration phenomena and provide theoretically grounded, detectable criteria for identifying trivial recalibration. Experiments on standard benchmarks demonstrate that our approach significantly improves model diagnostic capability, yielding a more reliable and interpretable evaluation framework for calibration research.
This work addresses key challenges in post-hoc calibration—namely nonlinear miscalibration, poor scalability to large numbers of classes, and perturbation of original predictions—by proposing Invertible Logit Transformation (InvLT). InvLT applies a shared-parameter scalar MLP element-wise to pre-softmax logits and incorporates a paired inverse network with soft monotonicity constraints. This design achieves high expressiveness and strong class scalability without introducing class-dependent parameters or requiring model retraining, while rigorously preserving the original classification accuracy. Extensive experiments across diverse image classification benchmarks and model architectures demonstrate that InvLT consistently outperforms existing calibration methods on standard calibration metrics, all while maintaining the original predictive performance without degradation.
Modern statistical models are growing increasingly complex in an effort to realistically capture system dynamics. Using standard simulation-based inference, these models may be computationally prohibitive, necessitating the use of model calibration methods. Bayesian score calibration is a computationally efficient framework for model calibration with strong theoretical guarantees. This framework learns an appropriate correction for an approximate model using a small number of simulations from the data-generating process. Currently, only a location-scale transformation has been explored, which may lack the flexibility to correct the complex error introduced by some approximate models. In this paper, we develop two flexible transformations for use in the Bayesian score calibration framework. The first is a polynomial extension, which can appropriately adjust approximate models with location-varying error. The second is a sequential application of Bayesian score calibration, which can accommodate approximate models with posteriors that have low support for the true parameter values. We also discuss an additional diagnostic for use with this framework. We demonstrate the increased flexibility these two approaches provide over Bayesian score calibration in two illustrative simulation studies.
This work addresses the poor calibration of deep neural networks, which often leads to substantial errors in low-confidence regions. Conventional temperature scaling struggles to correct heterogeneous miscalibration across the confidence spectrum due to its global, uniform adjustment. To overcome this limitation, the authors propose a quantile-aware adaptive temperature scaling method that maps predicted confidences into quantile space and constructs a monotonic temperature function adaptively varying with empirical confidence quantiles. This approach explicitly models confidence heterogeneity and naturally aligns with a reparameterized expected calibration error (ECE) objective. Experimental results demonstrate that the method consistently outperforms existing post-hoc calibration techniques across diverse datasets, model architectures, and distribution shift scenarios, yielding more reliable confidence estimates without altering the original predictions.
Existing methods struggle to simultaneously achieve high accuracy, robustness, and calibration in neural networks. This work proposes Lipschitz Scaling Training (LiST), which establishes, for the first time, a theoretical connection between Lipschitz constraints and temperature scaling. By dynamically adjusting the global Lipschitz constant during training, LiST embeds calibration directly into the learning process, automatically identifying a calibration-optimal operating point along the accuracy–robustness Pareto frontier. The method integrates margin-aware Lipschitz constraints, dynamic constant adaptation, and calibration-aware optimization, and further improves sample efficiency by reusing calibration data after convergence. Experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet demonstrate that LiST matches baseline performance in both accuracy and robustness while achieving well-calibrated predictions without any post-hoc processing.
This study addresses the challenge of jointly modeling calibration and control parameters in computer model calibration, where the distribution of calibration parameters is unknown while that of control parameters is known. To tackle this issue, the authors propose a nonparametric Bayesian calibration method based on measure decomposition. The approach preserves the known marginal distribution of the control parameters while employing stochastic process modeling and Bayesian inference to construct a posterior distribution over the input space that aligns with field observations. Notably, this work is the first within a nonparametric calibration framework to explicitly maintain the prior distributional properties of the control parameters, thereby substantially enhancing the physical consistency and scientific credibility of the calibration results.
This work addresses the longstanding challenge in calibration evaluation: existing metrics struggle to simultaneously achieve operability—formally tied to full swap regret—and testability, which requires accurate estimation from limited samples. To bridge this gap, we propose the Soft-binned Calibration Decision Loss (SCDL), a novel calibration measure constructed via a soft-binning strategy that yields a continuous and statistically consistent estimator. Leveraging tools from probabilistic calibration theory, swap regret analysis, and statistical estimation, we establish that SCDL is the first metric to theoretically guarantee full operability while achieving near-optimal sample complexity for testability. Empirical evaluations further demonstrate that SCDL substantially outperforms current state-of-the-art methods in practical calibration assessment.