🤖 AI Summary
This work addresses conformal prediction under the challenging setting of *unlabeled calibration data*, proposing the first general framework that achieves statistically valid coverage guarantees without requiring any labeled samples. Methodologically, it leverages unlabeled data to estimate a surrogate calibration score and derives a verifiable coverage bound by incorporating the model’s accuracy (or a precision metric). Theoretically, the resulting prediction set satisfies a finite-sample coverage guarantee: $mathbb{P}(Y in C) geq 1 - alpha - eta$. The framework is unified across both classification and regression tasks, eliminating the conventional reliance on labeled calibration sets. This breakthrough significantly broadens the applicability of conformal prediction to low-resource, privacy-sensitive, and high-label-cost scenarios. Moreover, it delivers a plug-and-play uncertainty quantification solution with rigorous statistical guarantees.
📝 Abstract
We extend the method of conformal prediction beyond the case relying on labeled calibration data. Replacing the calibration scores by suitable estimates, we identify conformity sets $C$ for classification and regression models that rely on unlabeled calibration data. Given a classification model with accuracy $1-β$, we prove that the conformity sets guarantee a coverage of $P(Y in C) geq 1-α-β$ for an arbitrary parameter $αin (0,1)$. The same coverage guarantee also holds for regression models, if we replace the accuracy by a similar exactness measure. Finally, we describe how to use the theoretical results in practice.