Improving measurement error and representativeness in nonprobability surveys

📅 2024-10-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Nonprobability surveys suffer from two long-overlooked sources of bias—measurement error and selection bias—neither of which is adequately addressed by conventional approaches. Method: We propose a composite estimator that jointly integrates probability and nonprobability samples, systematically correcting for both measurement error (e.g., via calibration or proxy modeling) and selection bias (e.g., via weighting or propensity scoring) within a unified finite-population mean estimation framework. Contribution/Results: Theoretically, our estimator achieves lower mean squared error than standard selection-bias-only corrections. Empirically, applied to U.S. business online survey data, it yields significantly higher accuracy than pure probability-survey estimators and aligns more closely with large-scale government benchmark surveys. This work establishes a novel, theoretically rigorous yet practically implementable pathway for scientifically leveraging nonprobability data in official and academic statistics.

Technology Category

Search and Optimization: Non-convex OptimizationReasoning under Uncertainty: Probabilistic InferenceMachine Learning: Calibration & Uncertainty Quantification

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSocial Networks and Social Media: Fairness and bias in social network and social media analysisSecurity and Privacy: Large-scale security measurements
📝 Abstract
Nonprobability surveys suffer from representativeness issues due to their unknown selection mechanism. Recent research on nonprobability surveys has primarily focused on reducing such selection bias. But bias due to measurement error is also present, as pointed out by Kennedy, Mercer and Lau (2024) using a benchmarking study in the case of commercial online nonprobability surveys in the United States. Before this study, measurement error bias in nonprobability surveys has mostly been overlooked and statistical methods have been devised for reducing only the selection bias, under the assumption of accuracy of survey responses. Motivated by this case study, our research focuses on combining two key areas in nonprobability sampling research: representativeness and measurement error, specifically aiming to mitigate bias from both sampling and measurement errors. In the context of finite population mean estimation, we propose a new composite estimator that integrates both probability and nonprobability surveys, promising improved results compared to benchmark values from large government surveys. Its performance in comparison to an existing composite estimator is analyzed in terms of mean squared error, analytically and empirically. In the context of the aforementioned case study, we further investigate when the proposed composite estimator outperforms estimator from probability surveys alone.
Problem

Research questions and friction points this paper is trying to address.

Addressing measurement and sampling errors in nonprobability surveys
Proposing a data integration method using multiple surveys
Leveraging machine learning to construct a composite estimator
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates multiple probability and nonprobability survey data
Leverages machine learning to construct composite estimator
Addresses both measurement and sampling errors simultaneously
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
University of Maryland, College Park
A
Aditi Sen
Applied Mathematics & Statistics, and Scientific Computation, University of Maryland, College Park, MD 20742, USA
P
Partha Lahiri
Department of Mathematics and The Joint Program in Survey Methodology, University of Maryland, College Park, MD 20742, USA