Score
Designs and evaluates calibration methods that adjust model scores by measuring an input’s distance or standardized deviation from the nearest class or prototype vectors and using that deviation to correct or rescale likelihood/novelty/outlier scores. This work covers computing standardized nearest‑prototype deviations in (often frozen) latent spaces, enforcing prototype consistency across examples, and combining prototype‑based deviation terms with existing scores (for example adding deviation to a standardized flow score) to correct ranking biases such as those caused by multimodal normals.
Existing studies lack rigorous theoretical characterization of posterior calibration methods—such as Platt scaling and isotonic regression—particularly regarding their dependence on feature quality, generalizability across models and datasets, and convergence behavior and robustness under finite-sample regimes. Method: We establish a unified theoretical framework for these two dominant calibration paradigms, deriving the first non-asymptotic guarantees on convergence rates, computational complexity, and explicit sample-size dependencies. Our analysis quantifies the relationship between feature informativeness and calibration robustness. Results: Through synthetic experiments and extensive empirical evaluation across diverse model architectures and benchmark datasets, we validate our theoretical findings. The results yield actionable guidance: isotonic regression is preferable under low signal-to-noise ratios or limited samples, whereas Platt scaling exhibits superior robustness in high-dimensional sparse feature settings. Our work provides interpretable, reusable principles for uncertainty calibration in practical machine learning systems.
This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.
This paper addresses the problem that conventional calibration evaluation of deep learning models is vulnerable to spurious recalibration—i.e., trivial post-hoc adjustments that improve calibration metrics without enhancing generalization. To tackle this, we propose a novel joint evaluation paradigm integrating calibration and generalization. First, we derive a Bregman-divergence-based decomposition of calibration error, establishing the first theoretical connection between calibration metrics and generalization objectives (e.g., negative log-likelihood). Second, we design a new reliability diagram that jointly visualizes calibration bias and estimated generalization error. Third, we characterize multiple “pseudo-optimal” calibration phenomena and provide theoretically grounded, detectable criteria for identifying trivial recalibration. Experiments on standard benchmarks demonstrate that our approach significantly improves model diagnostic capability, yielding a more reliable and interpretable evaluation framework for calibration research.
Normalized flow generative models suffer from interpolation paths deviating from the data manifold, primarily due to norm drift induced by Gaussian base distributions in latent space. To address this, we propose a norm-constrained base distribution reconstruction framework—introducing Dirichlet and von Mises–Fisher distributions into normalized flows for the first time. These distributions explicitly constrain latent variables to the unit simplex or unit hypersphere, respectively, ensuring geometrically consistent interpolation trajectories. Our method requires no architectural modifications to the flow network and provides an interpretable, unambiguous interpolation criterion, effectively overcoming interpolation distortion inherent to the Gaussian assumption. Experiments demonstrate consistent improvements over baselines across all major evaluation metrics: bits/dim, Fréchet Inception Distance (FID), and Kernel Inception Distance (KID). Interpolation quality is significantly enhanced while strictly preserving original generation performance.
Existing calibration error estimation lacks differentiable, optimizable estimators, hindering end-to-end calibration optimization. Method: We formulate the squared calibration error estimation as a regression task over i.i.d. sample pairs, adopting mean-squared error (MSE) as the risk criterion. Leveraging the bilinear structure of the squared calibration error, we employ kernel ridge regression with joint hyperparameter optimization within a novel train-validation-test estimation pipeline. Contribution/Results: This work establishes the first unified risk-based framework for calibration error estimation; reformulates canonical calibration error estimation as a learnable, differentiable regression problem; and introduces a principled three-stage estimation protocol. Evaluated on standard image classification benchmarks, our estimator achieves significantly higher accuracy than state-of-the-art methods. It is the first practical, end-to-end optimizable estimator for canonical calibration error, enabling gradient-based calibration refinement.
This work addresses the challenge in multi-class anomaly detection where unified models often suffer from anomaly replication and confusion among normal classes. To this end, we propose a label-free training framework that formulates the task as a representational capacity allocation problem. By leveraging a shared learnable prototype bank, our approach introduces a dual regularization mechanism—spatial prototype alignment and prototype-relative global alignment—to enhance reconstruction fidelity for normal samples and suppress anomaly replication, all without requiring class labels, negative samples, or memory-based retrieval. The method preserves the standard teacher–student feature discrepancy pipeline while significantly improving both the separation between anomaly and normal scores and the discriminability among normal categories. It achieves state-of-the-art average detection accuracies of 86.2%, 80.7%, and 73.1% on MVTec AD, VisA, and Real-IAD benchmarks, respectively.
This study addresses the unreliability of probability estimates from modern classifiers and the absence of a unified, large-scale evaluation framework for post-hoc calibration methods. The authors construct the first comprehensive calibration benchmark encompassing nearly 2,000 experiments across tabular and computer vision tasks, integrating classical models, deep networks, and foundation models, and systematically reimplement dozens of calibration techniques within a consistent framework. They introduce a novel metric, Post-hoc Improvement (PHI), which combines proper scoring rules to jointly assess calibration quality and predictive performance. Key findings reveal that smoothing-based calibration consistently outperforms binning approaches, high-dimensional multiclass settings demand specialized strategies, and off-the-shelf foundation models exhibit poor calibration without explicit design considerations. All data, code, and tools are publicly released to enable plug-and-play research.
Existing pose flow–based anomaly detection methods rely on a single flow score, which struggles to capture the multimodality of normal behaviors and is sensitive to pose observation noise, further lacking an effective calibration mechanism under frozen detector settings. This work proposes a lightweight post-processing calibration approach that, for the first time, integrates nearest prototype deviation in latent space with keypoint confidence gating to enable reliability-aware recalibration of the original flow scores—without requiring model retraining. Evaluated across two backbone networks and four benchmark datasets, the method consistently improves frame-level AUROC by 0.34–4.49 percentage points, achieving an average gain of 2.03 percentage points.
This study addresses the challenge of multi-source evaluation in the absence of ground-truth labels and a shared annotation space, where incomparable output scales across scorers hinder the construction of effective supervision signals. To overcome this, we propose a calibration-first framework that synthesizes a universal ordinal reference space via ordered calibration features, aligning subset-specific scorers to a unified scale. This approach is further augmented by a low-resolution calibration approximation technique, enabling supervision score fusion independent of training distributions. Evaluated on three benchmark datasets, the proposed method significantly outperforms both uncalibrated averaging and the best individual scorer while substantially reducing computational costs, thereby demonstrating its effectiveness and generalizability in ground-truth-free scenarios.
This work addresses the susceptibility of existing calibration evaluations for large language models to accuracy disparities, which distorts cross-model comparisons. To enable fair calibration assessment while controlling for accuracy, the authors propose the ACE framework, incorporating three alignment mechanisms: instance alignment, distribution alignment, and candidate alignment. The study systematically reveals, for the first time, the bias inherent in conventional global calibration metrics—such as Expected Calibration Error and Brier Score—when used for comparing models with differing accuracies, and introduces an accuracy-controlled correction strategy. Experimental results demonstrate that the apparent calibration advantages of most models substantially diminish—and their rankings frequently reverse—once calibration metrics are adjusted for accuracy, thereby demonstrating that unadjusted metrics are unsuitable for cross-model calibration evaluation.