🤖 AI Summary
This study addresses the coarsening of self-reported numeric variables in surveys—often caused by rounding or heaping—by proposing a novel approach that integrates design-based inference with latent variable modeling. Treating observed values as coarsened manifestations of an underlying continuous latent variable, the method jointly models the coarsening mechanism and the latent distribution via a survey-weighted pseudo-likelihood. It generates posterior predictive replicates to propagate coarsening-induced uncertainty into standard design-based estimators. This framework is the first to explicitly correct for coarsening bias under complex sampling designs, enabling unbiased estimation of means, quantiles, and threshold-based prevalence measures. Simulation studies demonstrate robustness across various model misspecifications and sampling scenarios, and empirical application to Italy’s PASSI behavioral surveillance data shows effective correction of coarsening-related estimation bias.
📝 Abstract
Self-reported numerical variables in sample surveys are frequently subject to coarsening, as respondents tend to report rounded or heaped values rather than the exact underlying quantity. While the statistical literature offers various model-based solutions, official statistics and public health surveillance routinely require design-based inference on finite-population indicators under complex sampling designs. This paper introduces a general framework that bridges this gap by treating the observed response as a coarsened manifestation of a latent variable, modeling reporting regimes of increasing coarseness jointly with the latent distribution via a survey-weighted pseudo-likelihood approach. The fitted model is used to generate posterior predictive replicates of the latent values, to which standard design-based estimators are applied, formally propagating the additional uncertainty through an explicit variance decomposition. Simulation studies assess the finite-sample performance and robustness of the proposed method under various misspecification and sampling scenarios. Finally, its practical relevance is demonstrated through an application to behavioural data from the Italian PASSI surveillance system, illustrating how coarsening affects different functionals, such as means, quantiles, and threshold-based prevalence indicators in heterogeneous ways.