🤖 AI Summary
This study addresses the exponential growth in calibration costs caused by likelihood ratios in weighted conformal prediction under covariate shift. To mitigate this issue, we propose a sketched calibration method that computes weights via compressed covariates, thereby reducing sensitivity to irrelevant shift directions and effectively controlling calibration costs. Theoretically, we prove that the compression operation does not increase shift-dependent calibration overhead and quantify response distribution leakage as a key metric for coverage guarantees. Empirically, our experiments demonstrate that the proposed approach reduces the proportion of infinite prediction sets from 37.8% to 8.0%, significantly improving predictive efficiency while maintaining the target coverage rate.
📝 Abstract
Weighted conformal prediction corrects for covariate shift by reweighting calibration scores with the likelihood ratio between target and source covariates. Its cost grows with the chi-square divergence between the two covariate laws, typically exponentially in the size of the shift, and much of it can be paid for shift in directions that do not affect the response. We propose sketched calibration: weighted conformal prediction with the ratio of compressed covariates $Z=T(X)$, used in the weights and optionally in the score. Compression never increases the shift-dependent calibration cost, and target coverage is at least $1-α-Δ_T$, where the leakage $Δ_T$ measures how much of the discarded shift reappears as a change in the law of the response, or of the score, given $Z$. The leakage is a covariance between the discarded shift and the response on the fibres of the sketch; it vanishes when the sketch is sufficient or retains the shift, and a minimax construction shows that no threshold rule for a fixed score can avoid it. In a heteroscedastic simulation where unweighted calibration fails, a one-dimensional sketch cuts the fraction of infinite prediction sets from $37.8\%$ to $8.0\%$ at mean coverage $92.3\%$ for a $90\%$ target; nonlinear and real-data experiments show when learned sketches succeed and when they fail.