PPI is the Difference Estimator: Recognizing the Survey Sampling Roots of Prediction-Powered Inference
This study investigates how to achieve valid statistical inference when combining machine learning predictions with a small number of gold-standard labels, and clarifies its connections to classical survey sampling methods. Through theoretical analysis, it establishes for the first time the algebraic equivalence between the core estimator in prediction-powered inference (PPI) and model-assisted estimators from the 1970s—such as difference and generalized regression (GREG) estimators. The work systematically compares these approaches in terms of inferential paradigms, use of unlabeled data, and subgroup estimation error, delineating which aspects of PPI are inherited versus novel. It further proposes directions for integrating insights from both fields. These results ground PPI in classical survey sampling theory while simultaneously expanding the toolkit available for modern, nonstandard estimators within the survey sampling framework.