🤖 AI Summary
This paper addresses the inefficiency (i.e., prediction set size) and generalization of weighted conformal risk control (W-CRC) under covariate shift. Methodologically, it integrates importance weighting, conformal prediction, and statistical learning theory to derive a computable upper bound on inefficiency during training—explicitly linking it to the base predictor’s generalization error, the degree of distribution shift, and sample size. The key contribution is the first quantitative characterization of how W-CRC efficiency depends on the base predictor’s generalization capability, and the theoretical revelation that prediction set informativeness decays with increasing shift severity. Empirical validation on fingerprint-based indoor localization demonstrates that the derived bound accurately captures the trend of prediction set size under varying shift levels, confirming both theoretical soundness and practical utility.
📝 Abstract
Predictive models are often required to produce reliable predictions under statistical conditions that are not matched to the training data. A common type of training-testing mismatch is covariate shift, where the conditional distribution of the target variable given the input features remains fixed, while the marginal distribution of the inputs changes. Weighted conformal risk control (W-CRC) uses data collected during the training phase to convert point predictions into prediction sets with valid risk guarantees at test time despite the presence of a covariate shift. However, while W-CRC provides statistical reliability, its efficiency -- measured by the size of the prediction sets -- can only be assessed at test time. In this work, we relate the generalization properties of the base predictor to the efficiency of W-CRC under covariate shifts. Specifically, we derive a bound on the inefficiency of the W-CRC predictor that depends on algorithmic hyperparameters and task-specific quantities available at training time. This bound offers insights on relationships between the informativeness of the prediction sets, the extent of the covariate shift, and the size of the calibration and training sets. Experiments on fingerprinting-based localization validate the theoretical results.