nearest-prototype calibration

Designs and evaluates calibration methods that adjust model scores by measuring an input’s distance or standardized deviation from the nearest class or prototype vectors and using that deviation to correct or rescale likelihood/novelty/outlier scores. This work covers computing standardized nearest‑prototype deviations in (often frozen) latent spaces, enforcing prototype consistency across examples, and combining prototype‑based deviation terms with existing scores (for example adding deviation to a standardized flow score) to correct ranking biases such as those caused by multimodal normals.

nearest-prototypecalibration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Calibration Meets Reality: Making Machine Learning Predictions Trustworthy

Sep 28, 2025
KP
Kristina P. Sinaga
🏛️ Independent Researcher

Existing studies lack rigorous theoretical characterization of posterior calibration methods—such as Platt scaling and isotonic regression—particularly regarding their dependence on feature quality, generalizability across models and datasets, and convergence behavior and robustness under finite-sample regimes. Method: We establish a unified theoretical framework for these two dominant calibration paradigms, deriving the first non-asymptotic guarantees on convergence rates, computational complexity, and explicit sample-size dependencies. Our analysis quantifies the relationship between feature informativeness and calibration robustness. Results: Through synthetic experiments and extensive empirical evaluation across diverse model architectures and benchmark datasets, we validate our theoretical findings. The results yield actionable guidance: isotonic regression is preferable under low signal-to-noise ratios or limited samples, whereas Platt scaling exhibits superior robustness in high-dimensional sparse feature settings. Our work provides interpretable, reusable principles for uncertainty calibration in practical machine learning systems.

Investigating feature quality impact on calibration performanceProviding practical guidelines for calibration method selectionUnderstanding theoretical performance of post-hoc calibration methods

This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.

calibrationclassificationhierarchical relations

This paper addresses the problem that conventional calibration evaluation of deep learning models is vulnerable to spurious recalibration—i.e., trivial post-hoc adjustments that improve calibration metrics without enhancing generalization. To tackle this, we propose a novel joint evaluation paradigm integrating calibration and generalization. First, we derive a Bregman-divergence-based decomposition of calibration error, establishing the first theoretical connection between calibration metrics and generalization objectives (e.g., negative log-likelihood). Second, we design a new reliability diagram that jointly visualizes calibration bias and estimated generalization error. Third, we characterize multiple “pseudo-optimal” calibration phenomena and provide theoretically grounded, detectable criteria for identifying trivial recalibration. Experiments on standard benchmarks demonstrate that our approach significantly improves model diagnostic capability, yielding a more reliable and interpretable evaluation framework for calibration research.

Developing visualization for calibration and generalization errorProving relationship between full and confidence calibration errorReassessing calibration metrics in machine learning

Principled Interpolation in Normalizing Flows

Oct 22, 2020
SG
Samuel G. Fadel
🏛️ University of Campinas | Leuphana University | Norwegian University of Science and Technology

Normalized flow generative models suffer from interpolation paths deviating from the data manifold, primarily due to norm drift induced by Gaussian base distributions in latent space. To address this, we propose a norm-constrained base distribution reconstruction framework—introducing Dirichlet and von Mises–Fisher distributions into normalized flows for the first time. These distributions explicitly constrain latent variables to the unit simplex or unit hypersphere, respectively, ensuring geometrically consistent interpolation trajectories. Our method requires no architectural modifications to the flow network and provides an interpretable, unambiguous interpolation criterion, effectively overcoming interpolation distortion inherent to the Gaussian assumption. Experiments demonstrate consistent improvements over baselines across all major evaluation metrics: bits/dim, Fréchet Inception Distance (FID), and Kernel Inception Distance (KID). Interpolation quality is significantly enhanced while strictly preserving original generation performance.

Addressing side effects of linear interpolation pathsEnabling principled interpolation through base distribution changesImproving interpolation in normalizing flow generative models

Optimizing Estimators of Squared Calibration Errors in Classification

Oct 09, 2024
SG
Sebastian G. Gruber
🏛️ German Cancer Consortium (DKTK) | German Cancer Research Center (DKFZ) | Goethe University Frankfurt | Inria | Ecole Normale Supérieure | PSL Research University

Existing calibration error estimation lacks differentiable, optimizable estimators, hindering end-to-end calibration optimization. Method: We formulate the squared calibration error estimation as a regression task over i.i.d. sample pairs, adopting mean-squared error (MSE) as the risk criterion. Leveraging the bilinear structure of the squared calibration error, we employ kernel ridge regression with joint hyperparameter optimization within a novel train-validation-test estimation pipeline. Contribution/Results: This work establishes the first unified risk-based framework for calibration error estimation; reformulates canonical calibration error estimation as a learnable, differentiable regression problem; and introduces a principled three-stage estimation protocol. Evaluated on standard image classification benchmarks, our estimator achieves significantly higher accuracy than state-of-the-art methods. It is the first practical, end-to-end optimizable estimator for canonical calibration error, enabling gradient-based calibration refinement.

Improving classifier calibration trustworthinessOptimizing squared calibration error estimatorsSelecting and tuning calibration estimators

Latest Papers

What's happening recently
View more

This work addresses the challenge in multi-class anomaly detection where unified models often suffer from anomaly replication and confusion among normal classes. To this end, we propose a label-free training framework that formulates the task as a representational capacity allocation problem. By leveraging a shared learnable prototype bank, our approach introduces a dual regularization mechanism—spatial prototype alignment and prototype-relative global alignment—to enhance reconstruction fidelity for normal samples and suppress anomaly replication, all without requiring class labels, negative samples, or memory-based retrieval. The method preserves the standard teacher–student feature discrepancy pipeline while significantly improving both the separation between anomaly and normal scores and the discriminability among normal categories. It achieves state-of-the-art average detection accuracies of 86.2%, 80.7%, and 73.1% on MVTec AD, VisA, and Real-IAD benchmarks, respectively.

anomaly detectionmulti-classone-for-all

This study addresses the unreliability of probability estimates from modern classifiers and the absence of a unified, large-scale evaluation framework for post-hoc calibration methods. The authors construct the first comprehensive calibration benchmark encompassing nearly 2,000 experiments across tabular and computer vision tasks, integrating classical models, deep networks, and foundation models, and systematically reimplement dozens of calibration techniques within a consistent framework. They introduce a novel metric, Post-hoc Improvement (PHI), which combines proper scoring rules to jointly assess calibration quality and predictive performance. Key findings reveal that smoothing-based calibration consistently outperforms binning approaches, high-dimensional multiclass settings demand specialized strategies, and off-the-shelf foundation models exhibit poor calibration without explicit design considerations. All data, code, and tools are publicly released to enable plug-and-play research.

calibration benchmarkmodel calibrationpost-hoc calibration

Existing pose flow–based anomaly detection methods rely on a single flow score, which struggles to capture the multimodality of normal behaviors and is sensitive to pose observation noise, further lacking an effective calibration mechanism under frozen detector settings. This work proposes a lightweight post-processing calibration approach that, for the first time, integrates nearest prototype deviation in latent space with keypoint confidence gating to enable reliability-aware recalibration of the original flow scores—without requiring model retraining. Evaluated across two backbone networks and four benchmark datasets, the method consistently improves frame-level AUROC by 0.34–4.49 percentage points, achieving an average gain of 2.03 percentage points.

frozen detectorpose-flowprototype calibration

This study addresses the challenge of multi-source evaluation in the absence of ground-truth labels and a shared annotation space, where incomparable output scales across scorers hinder the construction of effective supervision signals. To overcome this, we propose a calibration-first framework that synthesizes a universal ordinal reference space via ordered calibration features, aligning subset-specific scorers to a unified scale. This approach is further augmented by a low-resolution calibration approximation technique, enabling supervision score fusion independent of training distributions. Evaluated on three benchmark datasets, the proposed method significantly outperforms both uncalibrated averaging and the best individual scorer while substantially reducing computational costs, thereby demonstrating its effectiveness and generalizability in ground-truth-free scenarios.

reference space calibrationscore fusionscorer alignment

This work addresses the susceptibility of existing calibration evaluations for large language models to accuracy disparities, which distorts cross-model comparisons. To enable fair calibration assessment while controlling for accuracy, the authors propose the ACE framework, incorporating three alignment mechanisms: instance alignment, distribution alignment, and candidate alignment. The study systematically reveals, for the first time, the bias inherent in conventional global calibration metrics—such as Expected Calibration Error and Brier Score—when used for comparing models with differing accuracies, and introduces an accuracy-controlled correction strategy. Experimental results demonstrate that the apparent calibration advantages of most models substantially diminish—and their rankings frequently reverse—once calibration metrics are adjusted for accuracy, thereby demonstrating that unadjusted metrics are unsuitable for cross-model calibration evaluation.

accuracy controlcalibrationevaluation metrics

Hot Scholars

XY

Xinge You

Professor of School of Electronics Information and Communications, Huazhong University of Science
Computer VisionPattern RecognitionMachine LearningWavelet Analysis and its Applications
GC

Guihai Chen

Professor of Computer Science
Computer Science and Technology
JC

Jiannong Cao

IEEE Fellow; Chair Professor, Hong Kong Polytechnic University
Distributed computingMobile and pervasive computingWireless sensor networksCloud computing
QP

Qinmu Peng

School of Electronics Information and Communications, Huazhong University of Science
Image processingPattern recognitionMedical image analysis