Score
Design, build, or analyze estimators that recover numeric calibration parameters and learnable transforms which map observed or degraded outputs back to nominal outputs—e.g., per‑channel multiplicative coefficients, additive biases, photometric or extrinsic transforms, or compact interpretable parameter vectors. These methods span statistical and regression formulations and include supervised, unsupervised/self‑supervised, few‑/k‑shot and data‑driven procedures that produce compact, interpretable, or task‑aware calibration estimates and can be applied jointly with task prediction without requiring full retraining.
Existing calibration error estimation lacks differentiable, optimizable estimators, hindering end-to-end calibration optimization. Method: We formulate the squared calibration error estimation as a regression task over i.i.d. sample pairs, adopting mean-squared error (MSE) as the risk criterion. Leveraging the bilinear structure of the squared calibration error, we employ kernel ridge regression with joint hyperparameter optimization within a novel train-validation-test estimation pipeline. Contribution/Results: This work establishes the first unified risk-based framework for calibration error estimation; reformulates canonical calibration error estimation as a learnable, differentiable regression problem; and introduces a principled three-stage estimation protocol. Evaluated on standard image classification benchmarks, our estimator achieves significantly higher accuracy than state-of-the-art methods. It is the first practical, end-to-end optimizable estimator for canonical calibration error, enabling gradient-based calibration refinement.
Existing post-hoc calibration methods for multi-class classifiers—particularly logistic regression–based approaches—suffer from overfitting due to excessive parameters and limited calibration data. Method: We propose Structured Matrix Scaling (SMS), a principled calibration framework built upon multinomial logistic regression, incorporating structured matrix regularization (e.g., low-rank or diagonal-plus-low-rank constraints), cross-class parameter sharing, and robust feature preprocessing. This design simultaneously enhances expressivity and controls variance. Contribution/Results: SMS theoretically overcomes the representational limitations of temperature scaling and standard matrix scaling, achieving a superior bias–variance trade-off. Extensive experiments across diverse models and datasets demonstrate that SMS significantly outperforms existing logistic regression–based calibration methods, while exhibiting strong scalability and practicality. The implementation is open-sourced, establishing a new efficient and robust benchmark for probabilistic calibration.
This paper investigates calibration of high-dimensional linear binary classifiers under the proportional asymptotic regime where the feature dimension $p$ and sample size $n$ grow at comparable rates ($p/n o gamma in (0,infty)$). We address miscalibration arising from bias in weight estimation inherent to conventional linear predictors. To resolve this, we propose *angular calibration*: a novel procedure that leverages a consistent estimator of the angle between the estimated and true weight vectors to interpolate between the learned classifier and a random-guess classifier, thereby achieving exact calibration. We establish that, in the high-dimensional limit, this method simultaneously satisfies both calibration and uniqueness-optimality under Bregman divergence—marking the first result to guarantee both properties jointly. Furthermore, we prove that Platt scaling converges to this optimal angular calibration solution in the proportional regime, providing a theoretical foundation for its empirical success.
This work addresses the suboptimal decoding decisions in large language models arising from the mismatch between predicted and true generation distributions. To mitigate this issue, the authors propose a task-aware calibration method that adjusts the model’s output distribution within a task-induced semantic latent space, integrated with Minimum Bayes Risk (MBR) decoding to yield improved decision-making. The study introduces, for the first time, calibration into task-specific latent semantic structures, establishing a new paradigm of task-aware calibration and proposing a task-oriented evaluation metric—Task Calibration Error (TCE). Experimental results demonstrate consistent improvements in generation quality across diverse tasks and model baselines, significantly enhancing the reliability of model decisions.
Existing studies lack rigorous theoretical characterization of posterior calibration methods—such as Platt scaling and isotonic regression—particularly regarding their dependence on feature quality, generalizability across models and datasets, and convergence behavior and robustness under finite-sample regimes. Method: We establish a unified theoretical framework for these two dominant calibration paradigms, deriving the first non-asymptotic guarantees on convergence rates, computational complexity, and explicit sample-size dependencies. Our analysis quantifies the relationship between feature informativeness and calibration robustness. Results: Through synthetic experiments and extensive empirical evaluation across diverse model architectures and benchmark datasets, we validate our theoretical findings. The results yield actionable guidance: isotonic regression is preferable under low signal-to-noise ratios or limited samples, whereas Platt scaling exhibits superior robustness in high-dimensional sparse feature settings. Our work provides interpretable, reusable principles for uncertainty calibration in practical machine learning systems.
Soft-label Bayesian error estimators are prone to severe distortion under non-genuine posteriors, potentially misestimating the irreducible error even when labels are calibrated. This work systematically analyzes the influence of temperature scaling on proxy measures of the Bayes error, providing for the first time an exact analytical expression for this distortion. We prove that the proxy value varies strictly monotonically with temperature and can be arbitrarily adjusted without altering classification performance. Under a Gaussian logits assumption, we derive a closed-form two-parameter solution. Experiments across eight binary classification tasks on CIFAR-10, Fashion-MNIST, and SVHN show that, while test error remains constant, the proxy value can vary by factors of 56 to 980; the closed-form estimates deviate from empirical values by less than 0.018, and the temperature minimizing expected calibration error does not correspond to a stable proxy estimate.
This work addresses the long-standing bottleneck of the $O(T^{2/3})$ lower bound on calibration error in online binary sequence calibration. The authors propose an efficient randomized predictor that integrates the SPR-Calibration procedure with an outer Blackwell-style correction mechanism, supported by a novel analytical framework based on proxy sequences and residual decomposition. By leveraging quadratic potential function analysis and exploiting sparsity structures, the method achieves—while maintaining computational efficiency—the first improvement over the classical bound, reducing the expected calibration error to $O(T^{2/3-\varepsilon})$ for some $\varepsilon > 0$, thereby significantly outperforming the previous best-known results.
This work proposes a ray-based camera calibration framework tailored for 3D reconstruction, addressing the limitations of traditional reprojection error–based methods that rely on 2D calibration boards and inadequately reflect 3D geometric accuracy. Instead of reprojection error, the approach introduces reconstruction error and intersection error as more representative metrics. It employs a novel icosahedral 3D calibration target and a ring-shaped feature detector, integrated with a generalized distortion model and bootstrapping to refine both intrinsic and extrinsic parameter estimates. Experimental results on synthetic data demonstrate that the proposed method reduces average intersection error by approximately 40%, significantly enhances calibration stability, and validates that ray-level metrics provide a more faithful assessment of 3D reconstruction fidelity compared to conventional approaches.
This work addresses key challenges in post-hoc calibration—namely nonlinear miscalibration, poor scalability to large numbers of classes, and perturbation of original predictions—by proposing Invertible Logit Transformation (InvLT). InvLT applies a shared-parameter scalar MLP element-wise to pre-softmax logits and incorporates a paired inverse network with soft monotonicity constraints. This design achieves high expressiveness and strong class scalability without introducing class-dependent parameters or requiring model retraining, while rigorously preserving the original classification accuracy. Extensive experiments across diverse image classification benchmarks and model architectures demonstrate that InvLT consistently outperforms existing calibration methods on standard calibration metrics, all while maintaining the original predictive performance without degradation.