Toward design-based inference for data integration

📅 2026-05-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unreliability of finite population inference from nonprobability samples, which often hinges on the unverifiable missing-at-random (MAR) assumption. The authors propose a design-based sequential sampling framework that treats the nonprobability sample as a deterministic stratum and draws a probability sample from its complement. Within this framework, they construct two classes of generalized regression estimators that guarantee design consistency without imposing any assumptions on the nonprobability selection mechanism—even under not-missing-at-random (NMAR) scenarios. Under stronger modeling assumptions such as coefficient homogeneity, the estimators achieve Isaki–Fuller asymptotic optimality. Simulations demonstrate that the proposed approach yields approximately unbiased estimates under both MAR and NMAR conditions and outperforms propensity score adjustment. Empirical analysis further reveals that separate regression is preferable when heterogeneity is strong, whereas combined regression offers modest efficiency gains under homogeneity.
📝 Abstract
Integrating non-probability samples into finite-population inference typically requires modeling unknown selection probabilities under a missing-at-random (MAR) assumption that is difficult to verify. We propose a design-based alternative in which the non-probability sample is treated as a fully observed certainty stratum and a probability sample is drawn only from the complementary, previously unsampled units. Within this sequential framework, we develop two generalized regression estimators: one fitting the outcome model separately in the complementary stratum, the other pooling both samples; we make two distinct contributions. First, both estimators are design-consistent and admit consistent variance estimators with no assumption whatsoever on the non-probability selection mechanism, including under not-missing-at-random (NMAR) selection. Second, under a working superpopulation model that holds in both strata, the pilot non-probability sample can be used to construct second-stage inclusion probabilities that achieve Isaki-Fuller asymptotic optimality for the separate estimator; this optimality claim relies on assumptions strictly stronger than MAR, but its failure does not invalidate the consistency results above. A diagnostic test for coefficient homogeneity is proposed to guide the choice between the two estimators. Simulations confirm that the sequential estimators remain essentially unbiased under both MAR and NMAR, while propensity-adjusted competitors can be severely biased under NMAR. Two applications from Lithuanian official statistics illustrate that separate regression is preferable when the pilot stratum and its complement are strongly heterogeneous, whereas combined regression offers a modest efficiency gain when the two strata are similar.
Problem

Research questions and friction points this paper is trying to address.

data integration
non-probability sampling
design-based inference
missing-not-at-random
finite-population inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

design-based inference
non-probability sampling
data integration
generalized regression estimator
Isaki-Fuller optimality
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Andrius Čiginas
Faculty of Mathematics and Informatics, Vilnius University, Vilnius, Lithuania
I
Ieva Burakauskaitė
Faculty of Mathematics and Informatics, Vilnius University, Vilnius, Lithuania
J
Jae Kwang Kim
Department of Statistics, Iowa State University, Ames, USA