🤖 AI Summary
This study addresses the challenge of insufficient conditional coverage in conformal prediction caused by scarce calibration data by proposing a prediction-driven quantile learning framework. The method leverages large language models (LLMs) to generate synthetic labels for augmenting the calibration set and reveals the geometric relationship between pinball risk and conditional coverage error. It further establishes a benefit-cost trade-off criterion for synthetic data and corrects biases through credible samples to optimize prediction set compactness. Experiments across eight benchmarks demonstrate that the proposed approach significantly improves conditional coverage while preserving marginal validity, yielding more compact prediction sets. These results validate the effectiveness of LLM-based annotation for enhancing conformal prediction under limited calibration data scenarios.
📝 Abstract
Conformal prediction provides distribution-free finite-sample marginal coverage, but post-hoc calibration data may be too scarce to learn how uncertainty varies across inputs. Meanwhile, abundant covariates can often be labeled cheaply by domain models or general-purpose language models. We study whether these synthetic labels can improve conditional coverage when only a small trusted sample is available. Building on score-quantile regression, we introduce prediction-powered quantile learning: a synthetic-labeled pool estimates pinball risk, paired trusted and synthetic outcomes correct its bias, and an independent trusted split performs final conformalization. Profiling pinball risk over scalar corrections reveals that population conditional-coverage error is its functional gradient; the corresponding Hessian removes global shifts and weights remaining shape error by boundary density. Composing this geometry with prediction-powered learning yields a three-resource expansion and a benefit--cost rule for synthetic power. Across eight regression benchmarks, synthetic-powered quantile learning substantially improves downstream conditional coverage while preserving marginal validity and producing more compact prediction sets. A human-rating study finds similar gains from external LLM labels and exposes a quality--quantity--cost tradeoff.