A comparison between geostatistical and machine learning models for spatio-temporal prediction of PM2.5 data

📅 2025-09-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional air quality monitoring networks suffer from insufficient spatiotemporal resolution, while low-cost PurpleAir sensor data exhibit systematic biases. Method: This study proposes a high-accuracy PM₂.₅ prediction framework integrating spatiotemporal dependency modeling and multi-algorithm ensemble learning. Conducting hourly modeling across California, we systematically benchmark kriging interpolation, land-use regression, neural networks, random forests, and support vector machines. We further introduce a novel spatiotemporally aware ensemble model that jointly enhances sensor data reliability and spatial heterogeneity representation via spatiotemporal graph convolution and bias correction. Contribution/Results: The proposed model significantly outperforms all individual baselines (RMSE reduced by 18.7%; R² increased by 0.12) and generates real-time, 1 km × 1 km resolution PM₂.₅ concentration maps—providing a scalable, cost-effective methodology for high-fidelity urban air quality monitoring.

Technology Category

Machine Learning: Ensemble MethodsPlanning, Routing, and Scheduling: Optimization of Spatio-temporal SystemsData Mining & Knowledge Management: Mining of Spatial, Temporal or Spatio-Temporal Data

Application Category

Graph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Web data quality in the era of algorithmically-generated content
📝 Abstract
Ambient air pollution poses significant health and environmental challenges. Exposure to high concentrations of PM$_{2.5}$ have been linked to increased respiratory and cardiovascular hospital admissions, more emergency department visits and deaths. Traditional air quality monitoring systems such as EPA-certified stations provide limited spatial and temporal data. The advent of low-cost sensors has dramatically improved the granularity of air quality data, enabling real-time, high-resolution monitoring. This study exploits the extensive data from PurpleAir sensors to assess and compare the effectiveness of various statistical and machine learning models in producing accurate hourly PM$_{2.5}$ maps across California. We evaluate traditional geostatistical methods, including kriging and land use regression, against advanced machine learning approaches such as neural networks, random forests, and support vector machines, as well as ensemble model. Our findings enhanced the predictive accuracy of PM2.5 concentration by correcting the bias in PurpleAir data with an ensemble model, which incorporating both spatiotemporal dependencies and machine learning models.
Problem

Research questions and friction points this paper is trying to address.

Comparing geostatistical and machine learning models for PM2.5 prediction
Addressing limited spatial-temporal data from traditional monitoring systems
Correcting bias in PurpleAir sensor data using ensemble modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Using ensemble model combining spatiotemporal dependencies
Correcting bias in low-cost sensor data
Comparing geostatistical and machine learning methods
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zeinab Mohamed
Department of Mathematics, Oberlin College and Conservatory
Wenlong Gong
Wenlong Gong
University of Houston System
Spatial statisticsBayesian hierarchical modelingGaussian processes