The risks of dichotomising ordinal outcomes: A spatial analysis of self-rated health in Western Europe

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the information loss and inferential bias arising from dichotomizing five-category self-rated health into a binary indicator. Utilizing European Social Survey data, we construct a Bayesian spatial cumulative logit model integrated with post-stratification techniques to systematically compare ordinal and binary regression frameworks, thereby quantifying how classification thresholds affect estimates of geographic health disparities across Western Europe. This work provides the first assessment of the sensitivity inherent in dichotomizing ordinal variables within spatial contexts, revealing that the allocation of the "fair" health category exerts a decisive influence on regional conclusions. While age–education gradients remain robust, geographic inferences prove highly sensitive to the chosen cut-point. Consequently, we recommend retaining full ordinal information or conducting rigorous sensitivity analyses when modeling spatial health variations.
📝 Abstract
Survey responses are often measured using ordered response categories. In the European Social Survey, self-rated health is measured on a five-point scale from very good to very bad, yet analyses commonly dichotomise responses into binary categories of "good" and "poor" health. Binary indicators provide prevalence measures that are straightforward to communicate, but dichotomisation reduces information and may limit captured health variation, while the cut-off may influence estimates and substantive conclusions. Systematic evidence on these effects in spatial settings remains limited. We address this gap using European Social Survey round 11 (2023/2024) for Western Europe. We model self-rated health among male respondents by age, education and region. Bayesian spatial individual-level models with poststratification provide population-representative estimates. We compare an ordinal cumulative logit model for the five-category outcome with Bernoulli logistic regression models using two dichotomisations differing in the classification of "fair" health. Older age and lower education are consistently associated with worse self-rated health across specifications, suggesting relatively robust fundamental age and educational gradients. However, their magnitude and uncertainty, and some geographical conclusions, are sensitive to how the outcome is modelled. Assigning "fair" health to either side of a binary cut-off changes the regions identified as having above-average levels of less favourable health. The ordinal model retains category-specific information and can produce familiar binary prevalence estimates through aggregation. Binary indicators remain useful, particularly for monitoring and communication. Nevertheless, the selected cut-off should be justified and sensitivity to alternative cut-offs or an ordinal modelling approach should be considered, especially for geographical comparisons.
Problem

Research questions and friction points this paper is trying to address.

dichotomisation
ordinal outcomes
self-rated health
spatial analysis
cut-off sensitivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bayesian spatial model
ordinal cumulative logit
dichotomisation
poststratification
self-rated health
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Miguel Ángel Beltrán-Sánchez
Department of Statistics and Operational Research, University of Valencia, Burjassot, Spain
M
Miguel Ángel Martínez-Beneito
Department of Statistics and Operational Research, University of Valencia, Burjassot, Spain
T
Terje Eikemo
Department of Sociology and Political Science, Norwegian University of Science and Technology, Trondheim, Norway
S
Sara Martino
Department of Mathematical Sciences, Norwegian University of Science and Technology, Trondheim, Norway
Andrea Riebler
Andrea Riebler
Department of Mathematical Sciences, NTNU, Trondheim, Norway
Statistics - Biostatistics