🤖 AI Summary
This study addresses the challenge that weighted split conformal prediction under covariate shift struggles to guarantee training-conditional coverage. To overcome this limitation, we derive an explicit training-conditional bound free of unspecified constants, revealing an intrinsic connection between the root-m convergence rate and chi-square divergence. Furthermore, we construct deterministic PAC prediction sets based on a variance-proxy scale and introduce a clipping strategy to optimize interval width. Our contributions include providing fully finite-sample theoretical guarantees and conducting a systematic comparison across multiple methods, thereby precisely delineating the conditions under which each approach yields narrower valid sets.
📝 Abstract
Weighted split conformal prediction reweights calibration scores by the likelihood ratio between the test and training covariate distributions and guarantees marginal coverage under covariate shift. We study its coverage conditional on the calibration data. An elementary argument, based on a single concentration inequality at a fixed population quantile, gives explicit training-conditional bounds without unspecified constants, and shows that the relevant scale is not the supremum of the likelihood ratio but a variance proxy built from the chi-squared divergence of the shift and from the average of the ratio over the part of the test population, of probability equal to the miscoverage level, where it is largest. A two-point lower bound shows that the root-m rate and the chi-squared contribution are intrinsic to the shift. Run at an explicitly inflated level, the weighted quantile becomes a deterministic PAC prediction set. We compare it with randomized rejection sampling and with importance-weighted learn-then-test and, through a certified choice of a clipping level for the likelihood ratio, map the regime in which each gives the narrower valid set. The analysis extends to estimated likelihood ratios and to tail functionals estimated from an unlabeled source sample, which yields a fully finite-sample certificate.