🤖 AI Summary
This study addresses the unreliability of finite population inference from nonprobability samples, which often hinges on the unverifiable missing-at-random (MAR) assumption. The authors propose a design-based sequential sampling framework that treats the nonprobability sample as a deterministic stratum and draws a probability sample from its complement. Within this framework, they construct two classes of generalized regression estimators that guarantee design consistency without imposing any assumptions on the nonprobability selection mechanism—even under not-missing-at-random (NMAR) scenarios. Under stronger modeling assumptions such as coefficient homogeneity, the estimators achieve Isaki–Fuller asymptotic optimality. Simulations demonstrate that the proposed approach yields approximately unbiased estimates under both MAR and NMAR conditions and outperforms propensity score adjustment. Empirical analysis further reveals that separate regression is preferable when heterogeneity is strong, whereas combined regression offers modest efficiency gains under homogeneity.
📝 Abstract
Integrating non-probability samples into finite-population inference typically requires modeling unknown selection probabilities under a missing-at-random (MAR) assumption that is difficult to verify. We propose a design-based alternative in which the non-probability sample is treated as a fully observed certainty stratum and a probability sample is drawn only from the complementary, previously unsampled units. Within this sequential framework, we develop two generalized regression estimators: one fitting the outcome model separately in the complementary stratum, the other pooling both samples; we make two distinct contributions. First, both estimators are design-consistent and admit consistent variance estimators with no assumption whatsoever on the non-probability selection mechanism, including under not-missing-at-random (NMAR) selection. Second, under a working superpopulation model that holds in both strata, the pilot non-probability sample can be used to construct second-stage inclusion probabilities that achieve Isaki-Fuller asymptotic optimality for the separate estimator; this optimality claim relies on assumptions strictly stronger than MAR, but its failure does not invalidate the consistency results above. A diagnostic test for coefficient homogeneity is proposed to guide the choice between the two estimators. Simulations confirm that the sequential estimators remain essentially unbiased under both MAR and NMAR, while propensity-adjusted competitors can be severely biased under NMAR. Two applications from Lithuanian official statistics illustrate that separate regression is preferable when the pilot stratum and its complement are strongly heterogeneous, whereas combined regression offers a modest efficiency gain when the two strata are similar.