Model Assisted Data Integration: An unbiased sampling strategy to use nonprobability data

📅 2026-04-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of achieving safe and efficient unbiased estimation using non-probability data in the absence of a controlled selection mechanism. The authors propose Model-Assisted Data Integration (MADI), a sampling strategy that integrates non-probability data with carefully designed probability samples and leverages arbitrary machine learning models to construct design-unbiased point estimators alongside corresponding unbiased variance estimators. MADI establishes, for the first time, a general framework for design-unbiased inference based on any machine learning model, offering both theoretical rigor and practical feasibility—particularly suited for high-frequency production environments in official statistics. Empirical results demonstrate that MADI substantially reduces estimation variance compared to traditional survey estimators, confirming its effectiveness and superiority.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Relational Probabilistic ModelsIntelligent Robots: State Estimation

Application Category

User Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingWeb Mining and Content Analysis: Web data integration and cleaningEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
The aim of survey statistics is to produce estimates with a minimal bias and a corresponding acceptable variance given a specific budget, preferable with a minor response burden for the participants. In recent years, considerable efforts have been made to achieve this through the extended use of found or non-probability data. However, to be able to safely utilize such data, rigorous theoretical foundations is needed, where one main concern is the of lack control due to not having access to the selection mechanism for the data. Several methods have been proposed in the literature to deal with this, though often relying on assumptions that may be difficult or impossible to verify in practice. Extending on the Data Integrated (DI) estimator introduced by Kim and Tam (2021), this paper introduce the Model Assisted Data Integration (MADI) sampling strategy. The proposed sampling strategy includes an estimator that has the desired properties: it is design-unbiased, has a design-unbiased variance estimator and is suitable for the intense production cycle of the statistical agency. The estimator uses nonprobability data combined with a probability sample that has a sampling design which aims to include individuals not captured by the nonprobability data. The estimator can use arbitrary machine learning models to produce unbiased estimates. A main conclusion of the paper is that the proposed sampling strategy can produce estimates with much lower variances compared to traditional survey estimators, and we use real empirical data to illustrate this point.
Problem

Research questions and friction points this paper is trying to address.

nonprobability data
survey estimation
selection mechanism
bias
variance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model Assisted Data Integration
nonprobability data
design-unbiased estimator
variance reduction
machine learning in survey statistics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Martin Hyllienmark
Statistics Sweden
G
Gustaf Strandell
Statistics Sweden