coverage evaluation

Designs, builds, and analyzes procedures and metrics that quantify and ensure the probability that model outputs (e.g., prediction sets, confidence intervals, or set-valued predictors) contain the true outcome — including estimating coverage probabilities, calibrating models to target coverage, measuring empirical coverage across axes, and designing coverage evaluation metrics and tests. Also develops methods to assess and maximize coverage under constraints (for example tradeoffs between risk and coverage), and to evaluate coverage behavior under domain differences or distributional shift.

coverageevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Conditional Coverage Diagnostics for Conformal Prediction

Dec 12, 2025
SB
Sacha Braun
🏛️ Inria Paris | Inria Paris-Saclay | UC Berkeley

Conditional coverage assessment remains a fundamental challenge in predictive reliability analysis, as existing conformal prediction methods only guarantee marginal coverage and lack localized (i.e., conditional) coverage guarantees. To address this, we propose a diagnostic framework centered on Excess Risk of Target coverage (ERT), which—novelty—formulates conditional coverage deviation as a learnable classification risk difference problem. Our approach leverages proper-loss-driven risk estimation, modern classifiers (e.g., deep neural networks), and theoretically grounded conservative upper bounds, enabling L1/L2-distance-based quantification, explicit separation of over- and under-coverage, and analysis under non-constant target coverage levels. Empirical evaluation demonstrates that ERT achieves significantly higher detection sensitivity than baselines such as CovGap. It has been successfully deployed for reliability benchmarking across diverse conformal methods and is accompanied by an open-source, unified evaluation toolkit.

Addresses sample inefficiency and overfitting in existing metricsBenchmarks conformal prediction methods with new ERT metricsEvaluates conditional coverage reliability in predictive systems

This study addresses the frequentist prohibition against assigning a probability to the coverage of a parameter by an observed confidence interval, which limits nuanced interpretation of coverage events. By embedding confidence interval construction within a unified probabilistic framework through thought experiments and formal modeling, the work introduces a coverage indicator variable and, from a multi-level conditional probability perspective, demonstrates the coherence of assigning intermediate probabilities to single-instance coverage events under specific regularity conditions. This approach transcends the strict behaviorist constraints traditionally imposed on confidence intervals, revealing a tension between the exclusive reliance on design-stage coverage probability and the definition of long-run error rates. The result is a more flexible and internally consistent theoretical foundation for interpreting confidence intervals.

confidence intervalscoverage probabilityfrequentist inference

Design-marginal calibration of Gaussian process predictive distributions: Bayesian and conformal approaches

Dec 05, 2025
AP
Aurélien Pion
🏛️ Transvalor S.A. | Univ. Paris-Saclay

This paper addresses the calibration of predictive distributions for Gaussian processes (GPs) under interpolation settings, formally defining μ-coverage and μ-probabilistic calibration via the randomized probability integral transform (RPIT) from a design-marginal perspective. We propose two novel methods: CPS-GP (Conformalized Predictive Smoothing GP), which achieves finite-sample marginal calibration, and BCR-GP (Bayesian-Constrained Residual GP), which yields smooth, sharp, and tail-controlled predictive distributions. Technically, both methods integrate leave-one-out residual standardization, generalized normal distribution modeling, cross-validated residual fitting, and Kolmogorov–Smirnov testing. Experiments demonstrate that CPS-GP and BCR-GP significantly outperform Jackknife+ and full-conformal GP in calibration metrics—including empirical coverage, KS statistic, and integrated absolute error—as well as in accuracy, measured by scaled continuous ranked probability score (CRPS). These advances provide a more reliable foundation for uncertainty quantification in applications such as sequential Bayesian optimization.

Calibrating Gaussian process predictive distributions for interpolationControlling dispersion and tail behavior in sequential design predictionsEnsuring marginal calibration through Bayesian and conformal methods

This work addresses a critical limitation of existing conformal prediction methods, which guarantee only marginal coverage and fail to characterize key operational metrics—such as decision frequency, error exposure, and rejection rate—and their inherent trade-offs in real-world deployment. To overcome this, the authors propose an operational certification framework that goes beyond coverage by introducing a calibration-audit two-stage mechanism to quantify and guarantee the statistical properties of system behavior under finite-sample settings. Key innovations include Small-Sample Beta Correction (SSBC) for finite-sample coverage guarantees, reusable confidence envelopes for operational metrics, and the revelation of geometric couplings and trade-off boundaries among these metrics under conformal partitioning. The framework successfully generates auditable operational configuration menus on Tox21 and AquaSolDB benchmarks, explicitly delineating performance boundaries and uncertainties across different calibration strategies.

conformal predictioncoveragedecision deferral

Bayesian sample size calculations for external validation studies of risk prediction models

Apr 22, 2025
MS
Mohsen Sadatsafavi
🏛️ the University of British Columbia | British Columbia Centre for Disease Control | Maastricht University | KU Leuven | University of Birmingham | National Institute for Health and Care Research

Conventional sample size calculations for external validation of risk prediction models rely on fixed prior assumptions about model performance (e.g., calibration, discrimination) and net benefit (NB), failing to capture real-world uncertainty; moreover, traditional precision-oriented approaches—based on confidence interval width—bear weak relevance to clinical utility (NB). Method: This paper introduces, for the first time, a systematic Bayesian framework for external validation sample size determination. It constructs a joint risk–outcome distribution from prior performance summaries and proposes a multi-objective Bayesian sample size rule that jointly optimizes expected precision, assurance probability, optimal NB identification, and expected value of sample information (EVSI). Contribution/Results: The method enables decision-making under quantified uncertainty, enhancing statistical robustness and clinical relevance. Applied to external validation of a COVID-19 deterioration risk model, it improves resource efficiency and reliability of validation conclusions.

Addressing limitations of conventional precision-based inference for net benefitBayesian sample size calculations for uncertain model performanceProposing rules for sample size based on assurance probabilities and EVSI

Latest Papers

What's happening recently
View more

This work addresses the limitation of traditional conformal prediction, which guarantees marginal coverage but often fails to achieve valid conditional coverage within subpopulations, while existing evaluation methods suffer from the curse of dimensionality. The authors propose a Conformal Prediction Analysis (CPA) framework that reframes conditional coverage assessment as a supervised learning task by training a reliability estimator to predict instance-level coverage probabilities. They introduce a Conditional Validity Index (CVI) to quantify the local safety and efficiency of conformal predictors. Theoretical analysis establishes the convergence of CVI and proves the consistency of CC-Select, a CVI-based model selection algorithm. Empirical results demonstrate that CPA effectively identifies local coverage failures and that CC-Select reliably selects models with superior conditional coverage.

Conditional CoverageConformal PredictionModel Selection

This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.

calibrationdistributional validationprobabilistic forecasting

This work addresses the limitations of existing conformal prediction methods, which typically guarantee only marginal coverage and struggle to ensure conditional coverage for heterogeneous test points or subpopulations, while lacking a unified theoretical framework to analyze their asymptotic validity, compare approaches, or extend them to structured data. The paper proposes the first unified theoretical framework tailored for conditional coverage, deriving non-asymptotic bounds on conditional miscoverage via pointwise and Lₚ paths. It systematically characterizes the sources of error underlying asymptotic conditional validity and provides a coherent interpretation of existing methods. Built upon a weighted symmetry formulation, the framework facilitates conditional coverage–oriented model selection, localization under covariate shift, and natural extensions to structured data. Numerical experiments corroborate the theoretical findings, establishing a comparable, extensible, and practically informative paradigm for conditional coverage.

asymptotic validityconditional coverageconformal prediction

This work addresses the deployment reliability challenges of current code language models, which often suffer from overconfidence or underconfidence due to the absence of effective uncertainty estimation and active abstention mechanisms. The authors propose a unified, deployment-oriented framework that treats uncertainty as an actionable signal, jointly optimizing model calibration, selective prediction, and lightweight program analysis tool invocation to establish an end-to-end decision-making and repair pipeline. Evaluated on both classification and generation tasks, the approach enables risk-controlled, coverage-adjustable applications, significantly improving correctness ranking and selective prediction performance while maintaining high coverage. This leads to a substantial reduction in error rates and enhances the practical reliability of code language models in real-world scenarios.

code language modelsmodel calibrationreliable code predictions

Traditional test adequacy metrics, such as code coverage and mutation testing, focus on implementation details and struggle to capture discrepancies between expected and actual program behavior. This work proposes an automated approach that extracts method-level expected behaviors from natural language documentation and source code, then maps them to existing test cases, thereby formalizing and empirically evaluating “behavioral gaps”—a dimension of test adequacy independent of structural metrics. By integrating natural language processing, static analysis, and behavioral mapping techniques, our method identifies 20,729 behaviors across ten Java open-source libraries with 93.1% precision, revealing that 17.5% of expected behaviors remain entirely untested. Notably, these gaps persist even in methods exhibiting high code coverage or high mutation kill rates, exposing a systematic deficiency in current testing practices—including automatically generated tests—in validating intended program behavior.

behavioural gapscode coverageexpected behaviour

Hot Scholars

OS

Osvaldo Simeone

King's College London
Information theorymachine learningquantum information processingwireless systems
SS

Shaojie Shen

Associate Professor, Hong Kong University of Science and Technology
Robotics
SD

Serge Demeyer

University of Antwerp
Software EngineeringSoftware EvolutionTest Automation
MI

Michael I. Jordan

Professor of Electrical Engineering and Computer Sciences and Professor of Statistics, UC Berkeley
machine learningcomputer sciencestatisticsartificial intelligence
NM

Nawshin Mannan Proma

Doctoral Researcher, University of York
Safety of AINavigation TechnologyGNSS Integrity Monitoring