calibration

Designs, implements, and evaluates methods and procedures to assess and adjust the agreement between a system’s outputs and observed ground truth, including calibration metrics, diagnostics, and recalibration algorithms. Builds tools and validation protocols to detect miscalibration across inputs or subpopulations and to apply corrective transformations or model retraining to improve the reliability of reported quantities.

calibration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This paper addresses the reliability of calibration evaluation for machine learning models, identifying systematic biases in the widely used Expected Calibration Error (ECE) under distributional shift and varying binning strategies. Methodologically, it clarifies the logical hierarchy among multi-level calibration definitions, and systematically exposes ECE’s limitations through visualization, binning-based statistical analysis, and theoretical derivation—demonstrating its failure to satisfy key requirements of robustness and consistency in calibration assessment. Building on this critique, the paper introduces and explicates emerging calibration paradigms—including distribution-level and instance-level calibration—alongside their corresponding evaluation methodologies, thereby constructing a rigorous, interpretable, and practice-oriented calibration knowledge framework. The results equip researchers with principled guidance for selecting appropriate evaluation metrics and advance calibration assessment from ad hoc, heuristic practices toward standardization and formalization.

Evaluation MetricsLimitationsMachine Learning Calibration

Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models

Aug 25, 2025
JW
Jelke Wibbeke
🏛️ Carl von Ossietzky Universität Oldenburg | German Aerospace Center (DLR) | Jade University of Applied Science

In safety-critical applications, evaluating uncertainty calibration of regression models is hindered by inconsistent metric definitions, conflicting assumptions, and incomparable scales—impeding interpretability and reproducibility. This work systematically surveys and categorizes existing calibration metrics, then conducts a model-agnostic benchmark across real-world, synthetic, and manually miscalibrated datasets. We empirically demonstrate—for the first time—that most metrics yield contradictory or even opposing conclusions for identical calibration states, confirming that metric choice critically influences research outcomes. To address this, we propose ENCE (Expected Normalized Calibration Error) and CWC (Weighted Coverage Confidence) as more robust and stable primary metrics. Experiments across diverse scenarios show that ENCE and CWC exhibit superior consistency, strong resilience to noise and distribution shifts, and enhanced interpretability. Our findings establish a reproducible methodological foundation for uncertainty calibration evaluation in regression.

Evaluating conflicting calibration metrics for regression modelsIdentifying inconsistencies in recalibration metric performance comparisonsSystematically benchmarking reliability of uncertainty quantification methods

This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.

calibrationdistributional validationprobabilistic forecasting

This paper addresses the problem that conventional calibration evaluation of deep learning models is vulnerable to spurious recalibration—i.e., trivial post-hoc adjustments that improve calibration metrics without enhancing generalization. To tackle this, we propose a novel joint evaluation paradigm integrating calibration and generalization. First, we derive a Bregman-divergence-based decomposition of calibration error, establishing the first theoretical connection between calibration metrics and generalization objectives (e.g., negative log-likelihood). Second, we design a new reliability diagram that jointly visualizes calibration bias and estimated generalization error. Third, we characterize multiple “pseudo-optimal” calibration phenomena and provide theoretically grounded, detectable criteria for identifying trivial recalibration. Experiments on standard benchmarks demonstrate that our approach significantly improves model diagnostic capability, yielding a more reliable and interpretable evaluation framework for calibration research.

Developing visualization for calibration and generalization errorProving relationship between full and confidence calibration errorReassessing calibration metrics in machine learning

Does In-IDE Calibration of Large Language Models work at Scale?

Oct 26, 2025
RK
Roham Koohestani
🏛️ Delft University of Technology | JetBrains Research | University of California, Davis

This study investigates the feasibility and effectiveness of confidence calibration for large language models (LLMs) in integrated development environments (IDEs), aiming to enhance the reliability of generated code and developer acceptance. Method: We propose a scalable, personalized calibration framework integrating post-hoc techniques (e.g., Platt scaling) and conduct a large-scale empirical analysis on 24 million real-world developer interaction logs, complemented by scenario-based design, semi-structured interviews, and questionnaire-based validation. Contribution/Results: General-purpose calibration fails to significantly improve confidence–accuracy alignment; in contrast, personalized calibration yields substantial gains for high-engagement users. Moreover, non-numerical, color-coded trust signals—designed specifically for IDE workflows—significantly improve usability and perceived trustworthiness. This work provides the first systematic characterization of boundary conditions and human-factor adaptation mechanisms for calibrating LLM-generated code suggestions.

Designing effective reliability communication for developersEvaluating calibration's impact on model confidence reliabilityInvestigating feasibility of calibrating code models in IDEs

Latest Papers

What's happening recently
View more

This work addresses the susceptibility of existing calibration evaluations for large language models to accuracy disparities, which distorts cross-model comparisons. To enable fair calibration assessment while controlling for accuracy, the authors propose the ACE framework, incorporating three alignment mechanisms: instance alignment, distribution alignment, and candidate alignment. The study systematically reveals, for the first time, the bias inherent in conventional global calibration metrics—such as Expected Calibration Error and Brier Score—when used for comparing models with differing accuracies, and introduces an accuracy-controlled correction strategy. Experimental results demonstrate that the apparent calibration advantages of most models substantially diminish—and their rankings frequently reverse—once calibration metrics are adjusted for accuracy, thereby demonstrating that unadjusted metrics are unsuitable for cross-model calibration evaluation.

accuracy controlcalibrationevaluation metrics

This work addresses the challenge that probabilistic programs generated by language models often suffer from statistical misspecifications—such as incorrect likelihoods, priors, or parameterizations—that are difficult to detect with conventional unit tests. The paper introduces, for the first time, Bayesian calibration as a central criterion for assessing the correctness of probabilistic programs and proposes a fully unsupervised, reference-free framework for their detection and repair. By integrating Bayesian validation techniques—including posterior predictive checks, simulation-based calibration (SBC), sampling diagnostics (e.g., $\hat{R}$, divergences, effective sample size), and held-out predictive log density—the method generates feedback signals to drive an iterative repair loop within large language models. Evaluated on 200 instances, the approach achieves detection AUCs of 0.97 with reference programs and 62–78% without, substantially outperforming unit testing; repair success rates reach 92% and 100% using GPT-5.1 and Claude, respectively.

Bayesian workflowcalibrationlanguage models

This work addresses the lack of fine-grained confidence calibration in large language models (LLMs) for automated code repair, which hinders developers’ ability to assess the reliability of model outputs. The study introduces fine-grained confidence calibration into this domain for the first time, applying localized Platt scaling to three distinct types of local edit confidence scores and integrating them with global calibration. This hybrid approach overcomes the limitations of conventional global calibration methods, which fail to capture uncertainty at the level of individual editing decisions. Extensive experiments across three code repair tasks and fourteen LLMs demonstrate that the proposed method significantly reduces calibration error, particularly improving calibration quality over a broader range of predicted probabilities, thereby enhancing the trustworthiness of model-generated code revisions.

automated code revisioncode repairconfidence calibration

This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.

distributional realignmenthigh-frequency monitoringintervention effect