conformal prediction

Constructing distribution-free, calibrated predictive sets or intervals (including online and group-conditional variants) that guarantee coverage properties for downstream decision tasks and can be integrated with surrogate models or active sampling.

conformalprediction

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Batch Predictive Inference

Sep 21, 2024
YL
Yonghoon Lee
🏛️ University of Pennsylvania

This paper addresses the problem of batch predictive inference (Batch PI) for multiple test points, proposing the first distribution-free framework that constructs prediction sets for unobserved sample functionals—such as the mean or median—with rigorous coverage guarantees, while simultaneously enabling conditional selection (e.g., false discovery rate control). Methodologically, it leverages exchangeability to design a rank-based quantile estimation scheme; for composable score functions (e.g., sample mean), it introduces a dedicated dynamic programming algorithm that avoids data splitting or naive conformal extensions. Theoretically, the method ensures exact, user-specified marginal coverage and maintains high informativeness even with small calibration sets. Experiments on synthetic benchmarks and a real-world drug–target interaction prediction task demonstrate substantial improvements over existing baselines in both coverage reliability and predictive efficiency.

Construct prediction sets for batch test pointsDevelop efficient algorithms for NP-hard quantile computationEnsure distribution-free coverage guarantees under exchangeability

Distribution-free inference with hierarchical data

Jun 10, 2023
YL
Yonghoon Lee
🏛️ University of Pennsylvania | University of Chicago

This work addresses the failure of standard exchangeability—and consequent unreliability of distribution-free inference—under hierarchical data structures (e.g., grouped or repeated-measures designs). We introduce *hierarchical exchangeability*, the first formal theoretical foundation for distribution-free inference in non-i.i.d. hierarchical settings. Methodologically, we extend conformal prediction and the jackknife+ framework to hierarchical structures and propose a second-moment coverage mechanism, strengthening guarantees from marginal coverage to *conditional second-moment coverage*. Experiments demonstrate that our approach substantially reduces conditional miscoverage rates; under model misspecification, prediction interval width increases only marginally, while under correct model specification, the overhead is negligible. Our core contribution is the first provably reliable, distribution-free inference framework tailored specifically for hierarchical data.

Achieves stronger second-moment coverage for repeated measurementsBalances prediction set width and conditional miscoverage ratesExtends distribution-free methods for hierarchical data structures

This paper addresses the conditional coverage performance of prediction intervals in regression. We propose a “conjecture testing” framework that imports hypothesis-testing principles into predictive inference; introduce— for the first time in nonparametric regression—the notion of *pertinence*, quantifying how well a prediction interval aligns with the local data distribution; and establish theoretically that model-free bootstrap achieves superior conditional coverage compared to quantile regression and substantially outperforms standard conformal prediction under mild regularity conditions. To enhance finite-sample reliability, we incorporate uniformization transformations and a refined conformal scoring function. Empirical results demonstrate significant improvements in interval pertinence under limited samples and support one-sided conjecture testing—thereby addressing key limitations of conformal prediction in both conditional coverage calibration and directional inference.

Compare Model-free Bootstrap and conformal prediction performanceEvaluate conditional coverage of prediction intervalsExtend pertinence concept to nonparametric regression

Confidence on the Focal: Conformal Prediction with Selection-Conditional Coverage

Mar 06, 2024
YJ
Ying Jin
🏛️ Harvard University | University of Pennsylvania

Under data-driven selection, conventional prediction intervals fail to guarantee marginal coverage for the selected units—compromising reliability for focal samples. Method: We propose the first finite-sample exact coverage framework for post-selection inference, extending Mondrian conformal prediction to multiple test samples and non-equivariant models while accommodating arbitrary permutation-invariant selection rules. Our approach integrates conditional randomization tests, top-K or optimization-driven selection, conformal p-values, and preliminary screening prediction sets to enable efficient computation. Contribution/Results: Evaluated on drug discovery and health risk prediction tasks, our method substantially improves empirical coverage for focal units, ensuring statistically valid inference in real-world decision-making scenarios. This provides the first provably exact finite-sample coverage guarantee for post-selection prediction intervals under general selection mechanisms.

Addresses selection bias in conformal predictionEnsures valid coverage for selected focal unitsGeneralizes to multiple test units and classifiers

Joint Coverage Regions: Simultaneous Confidence and Prediction Sets

Mar 01, 2023
ED
Edgar Dobriban
🏛️ University of Pennsylvania

This paper addresses the challenge of simultaneously covering an unknown fixed parameter (e.g., the population mean) and an unobserved random data point (e.g., a new test sample) within the frequentist framework. We propose joint coverage regions (JCRs)—the first method unifying confidence sets and prediction sets under frequentist inference. JCRs are constructed via conditional pivotal quantities, leveraging data splitting and set optimization to guarantee finite-sample joint coverage probability for both the parameter and the new observation. Unlike conventional separate inference procedures, JCRs achieve both statistical validity and computational tractability. Empirical evaluations on mean estimation and structured prediction tasks demonstrate substantial improvements in set compactness and practical utility. Crucially, JCRs establish the first theoretical unification of classical confidence inference with distribution-free prediction, bridging two foundational paradigms in statistical learning.

Cover unknown parameters and unobserved datapoints simultaneouslyDevelop efficient algorithms for Joint Coverage RegionsUnify confidence intervals and prediction regions in statistics

Latest Papers

What's happening recently
View more

Existing online conformal prediction methods struggle to simultaneously achieve parameter-free operation and rigorous error control under group-conditional settings in non-stationary data streams, limiting the fairness and robustness of uncertainty quantification. This work proposes the first online conformal prediction algorithm that unifies parameter-free online optimization with group-conditional coverage guarantees. The method requires no learning rate tuning and provides strict conditional coverage for distinct data groups even in dynamic environments. Empirical evaluations demonstrate that the proposed approach substantially improves the reliability of current parameter-free methods while delivering prediction interval quality comparable to carefully tuned group-conditional baselines.

distribution shiftgroup-conditional coverageonline conformal prediction

This work addresses a key limitation of traditional conformal prediction, which guarantees only marginal coverage without independent control over upper and lower tail risks. The authors propose a novel split-conformal approach that constructs one-sided prediction intervals—each marginally valid for its respective tail—and derives a two-sided interval via their intersection. This framework provides, for the first time, finite-sample or asymptotic guarantees of tail-specific coverage. By moving beyond the conventional focus on overall coverage, the method significantly improves directional calibration in simulations with skewed data and enables practical applications such as financial portfolio optimization, where it effectively supports return maximization while rigorously controlling left-tail risk.

Conformal PredictionDirectional CalibrationMarginal Validity

This work addresses the selection bias inherent in selective inference when inference is performed only on high-informativeness prediction sets, which can compromise false coverage rate (FCR) control. Under an idealized setting, the authors derive an oracle strategy that maximizes statistical power while maintaining FCR guarantees. They then develop a finite-sample calibration mechanism grounded in conformal prediction and probability calibration to adapt this optimal strategy to practical settings. The resulting procedure rigorously controls FCR while achieving substantially higher power than existing methods. Empirical evaluations on both synthetic and real-world classification datasets demonstrate consistent superiority over current approaches in terms of statistical efficiency without sacrificing FCR control.

Conformal PredictionFalse Coverage RateInformative Prediction Sets

Existing predictive inference methods rely on assumptions of independent and identically distributed data or exchangeability, which fail to provide valid conditional coverage guarantees for structured data exhibiting group symmetries—such as networks or hierarchical clusters. This work proposes C-SymmPI, the first framework to achieve approximate conditional coverage for general structured data that are non-exchangeable yet invariant under distributional symmetries. The approach reformulates conditional coverage as miscoverage error over user-specified function classes and integrates relaxed multi-accuracy with invariance theory to devise a projection algorithm suitable for high-dimensional observations and sampling strategies for large or infinite symmetry groups. Experiments demonstrate that C-SymmPI substantially outperforms existing methods on hierarchical and network data, yielding prediction intervals that are more accurate, stable, and informative, while preserving theoretical guarantees under distributional shifts—thereby unifying and extending state-of-the-art results from the exchangeable setting.

conditional predictive inferencecoverage guaranteesdistribution-free inference

This work addresses the challenge that temporal dependence and cross-sectional heterogeneity in panel data violate the exchangeability assumption required by conformal prediction, leading to unreliable uncertainty quantification. To overcome this, the authors propose an online conformal prediction framework tailored for non-exchangeable panel data, featuring two key innovations: first, constructing calibration sets by weighting contemporaneous units based on historical similarity; second, adaptively adjusting miscoverage levels using sparse feedback. The method is distribution-free, compatible with any base predictive model, and demonstrably improves worst-case unit coverage on both synthetic and real-world datasets. By adaptively allocating prediction interval widths, it achieves more efficient and balanced uncertainty quantification across heterogeneous units.

conformal predictionnon-exchangeabilitypanel data

Hot Scholars

OS

Osvaldo Simeone

King's College London
Information theorymachine learningquantum information processingwireless systems
ZZ

Zhixin Zhou

Alpha Benito Research
StatisticsMachine Learning
MI

Michael I. Jordan

Professor of Electrical Engineering and Computer Sciences and Professor of Statistics, UC Berkeley
machine learningcomputer sciencestatisticsartificial intelligence
CZ

Changliang Zou

Professor of Statistics, Nankai University
StatisticsQuality Control
HW

Hongxin Wei

Southern University of Science and Technology (SUSTech)
Reliable Machine LearningUncertainty EstimationStatistics