🤖 AI Summary
This study addresses the challenge of accurate temporal forecasting of diabetes prevalence at the U.S. city level, tackling real-world issues including missing and inconsistent multi-source heterogeneous data and sparse ground-truth labels. We construct a standardized, comprehensive diabetes feature dataset covering major U.S. cities from 2011 to 2021. To this end, we propose EBMBag+, the first time-aware enhanced Bagging ensemble model: it unifies heterogeneous data sources via data-driven feature engineering and introduces time-weighted resampling alongside dynamic base-learner fusion—enhancing temporal generalization while preserving ensemble robustness. Evaluated on real-world data, EBMBag+ achieves MAE = 0.41, RMSE = 0.53, MAPE = 4.01%, and R² = 0.90, outperforming six baselines—including SVM, BDTree, LSBoost, NN, LSTM, and the benchmark ERMBag. The model delivers an interpretable, deployable predictive tool for precision public health interventions.
📝 Abstract
Diabetes is a chronic metabolic disease characterized by elevated blood glucose levels, leading to complications like heart disease, kidney failure, and nerve damage. Accurate state-level predictions are vital for effective healthcare planning and targeted interventions, but in many cases, data for necessary analyses are incomplete. This study begins with a data engineering process to integrate diabetes-related datasets from 2011 to 2021 to create a comprehensive feature set. We then introduce an enhanced bagging ensemble regression model (EBMBag+) for time series forecasting to predict diabetes prevalence across U.S. cities. Several baseline models, including SVMReg, BDTree, LSBoost, NN, LSTM, and ERMBag, were evaluated for comparison with our EBMBag+ algorithm. The experimental results demonstrate that EBMBag+ achieved the best performance, with an MAE of 0.41, RMSE of 0.53, MAPE of 4.01, and an R2 of 0.9.