🤖 AI Summary
This study addresses the limitation of conventional prediction-powered inference, which relies solely on point predictions and fails to exploit the full predictive distributions provided by models such as large language models, thereby constraining statistical efficiency. To overcome this, we propose DiPPI, a framework that systematically incorporates predictive distributions as auxiliary information into the inference process for the first time. Methodologically, DiPPI employs cross-fitted estimators and score calibration techniques to achieve distribution-aware modeling, while deriving optimal efficiency and theoretical guarantees under finite-dimensional representations. Simulation studies and real-data applications demonstrate that DiPPI significantly outperforms traditional point-prediction-based approaches, yielding substantial improvements in statistical efficiency.
📝 Abstract
Prediction-powered inference (PPI) typically relies on point predictions on unlabeled data. When predictive distributions are available as in a wide range of applications, including predictions from LLMs, we introduce distribution-informed prediction-powered inference (DiPPI), a general framework for further improving statistical efficiency by using predictive distributions as auxiliary information in the spirit of PPI. We characterize the optimal use of this information through score calibration, derive the oracle efficiency for a finite-dimensional representation of the predictive distribution, and provide theoretical guarantees for positive learning with the cross-fitted DiPPI estimator. Through simulations and three real-data applications, we show that DiPPI achieves better efficiency than PPI methods based on point predictions. These gains arise when the predictive distribution contains score-relevant information that is partially lost in point predictions.