Score
Designs, fits, and evaluates statistical and machine‑learning models for spatially referenced data that explicitly represent spatial dependence or autocorrelation across locations — including geostatistical covariance structures, spatial random effects and hierarchical/mixed models, spatially‑aware random forests, and spatial point‑process or zero‑inflated spatial models. These models are used to produce spatially informed predictions, quantify residual spatially structured variation or clusters, control for spatial dependency during feature selection, and estimate effects with reduced spatially structured error.
Existing surveys predominantly focus on single domains or model types, lacking interdisciplinary and systematic syntheses of spatiotemporal statistical models. Method: Guided by the PRISMA framework, we systematically identified and structurally analyzed 83 high-quality publications from 2021–2025 across epidemiology, ecology, public health, economics, and criminology. Contribution/Results: We propose a novel, unified classification framework for spatiotemporal models, revealing the ubiquity of hierarchical structures and additive spatiotemporal dependence. The analysis uncovers domain-specific modeling preferences and pronounced imbalances in research distribution. Critically, we find widespread deficiencies in reproducibility and methodological transparency—impeding cross-domain knowledge transfer. Our framework provides both a theoretical foundation and practical pathway for interdisciplinary model comparison, methodological borrowing, and principled model refinement.
This study addresses the lack of systematic comparisons among effective approaches for handling spatial autocorrelation in random forests. Integrating machine learning with geostatistical principles, it presents the first unified evaluation of four strategies—Gaussian process augmentation, observation-driven correlation structures, spatial basis functions, and local geographical fitting—to enhance the spatial predictive performance of random forests. Through both simulation experiments and an empirical analysis of air pollution in Blantyre, Malawi, the research demonstrates that while no single method consistently outperforms others across all scenarios, spatial basis functions exhibit robust and superior performance throughout. These findings offer practical and reliable guidance for spatial modeling of environmental processes using random forests.
Traditional geostatistical methods rely on second-order moments and Gaussian assumptions, which are inadequate for capturing non-Gaussian spatial dependence. This work addresses this limitation by leveraging Sklar’s theorem to introduce a copula-based framework that decouples marginal distributions from the spatial dependence structure. The authors systematically develop spatial copula models applicable at both fixed point sets and process levels, emphasizing Kolmogorov consistency to clarify distinctions between these two modeling paradigms. The framework is further extended to spatio-temporal settings and supports flexible marginal specifications. By integrating spatial statistics, copula theory, and stochastic processes, this study establishes a unified approach for modeling non-Gaussian spatial dependence, elucidating the relationships, strengths, and limitations of existing methodologies, thereby advancing both theoretical understanding and practical applications in the field.
Modeling large-scale, non-Gaussian, noisy, and incomplete spatial data remains challenging due to limitations of conventional copula models in capturing complex spatial heterogeneity and non-Gaussian dependence structures. Method: We propose a Bayesian hierarchical model integrating vine copulas with spatial random effects. Crucially, we design a novel vine copula structure explicitly embedding spatial dependence, enabling low-rank latent process representations and computationally efficient Bayesian inference. Contribution/Results: Our method overcomes key bottlenecks in traditional copula-based spatial modeling. In both parameter estimation and spatial prediction tasks, it significantly outperforms benchmarks—including fixed-rank kriging (FRK)—in accuracy, convergence speed, and robustness. Applied to atmospheric methane concentration mapping over Australia’s Bowen Basin using Sentinel-5P satellite remote sensing data, the approach delivers superior predictive performance under missingness and noise. This work establishes a new paradigm for non-Gaussian spatial statistical modeling.
This paper addresses spatial confounding under multivariate disease dependence, characterizing its dual interference mechanisms: from the modeling perspective, spatial random effects inflate posterior variances of fixed effects; from the data-generating perspective, correlations between covariates and unobserved spatial confounders induce biased inference. We innovatively distinguish variance inflation due to spatial confounding from that caused by multicollinearity, and propose the first diagnostic framework for spatial confounding in multivariate areal data. We theoretically prove—and empirically verify via large-scale simulation and U.S. county-level obesity, diabetes, and cancer mortality data—that incorporating spatial effects improves estimation efficiency even under spatial confounding and model misspecification. Leveraging Bayesian coregionalization, posterior variance decomposition, and hierarchical spatial modeling, our work establishes theoretical foundations and empirical evidence for the robustness of multivariate spatial models in complex, dependent disease settings.
Spatial confounding can induce bias in parameter estimation when modeling spatially structured data, thereby compromising the validity of statistical inference. This study presents the first unified framework that integrates approaches from both spatial statistics and causal inference for addressing spatial confounding. It systematically reviews the conceptual definitions, classical models, and cutting-edge strategies, and establishes a comprehensive comparative framework tailored to areal and geostatistical data. Through theoretical analysis and empirical evaluations across multiple real-world datasets, the work elucidates the performance disparities among existing methods under varying conditions, clarifies their respective applicability boundaries, and offers principled guidance for method selection and future research directions.
This study addresses the challenge that traditional spatial prediction models struggle to simultaneously maintain interpretability of covariates and flexibility in modeling complex spatial dependence structures. To overcome this limitation, the authors propose a semiparametric spatial autoregressive model that integrates linear covariate effects with nonparametrically estimated spatial components. This approach preserves model interpretability while flexibly capturing intricate spatial dependencies, thereby relaxing the strong assumptions on covariance structures commonly imposed by conventional models. The proposed method achieves both high predictive accuracy and strong interpretability, supported by a rigorous asymptotic theory. Empirical evaluations on both simulated and real-world datasets demonstrate that its predictive performance is comparable to that of geostatistical methods, while substantially outperforming classical spatial econometric models in terms of explanatory power.
This study addresses the unresolved question of when spatial and non-spatial random effects yield equivalent posterior inference for regression coefficients in Bayesian regression with multilevel areal data. Within a hierarchical Bayesian framework assuming Gaussian responses and employing a Leroux conditional autoregressive (CAR) prior, the authors formally derive—for the first time—a closed-form sample size threshold $m^*$ that determines whether spatial modeling is necessary. This threshold admits a clear interpretation, revealing that the difference in posterior variances converges to zero at an $O(m^{-1})$ rate. Simulation studies confirm that $m^*$ accurately identifies the tipping point in modeling complexity; notably, spatial modeling remains essential regardless of sample size whenever covariates exhibit no within-area variation.
This study addresses a key limitation of traditional marked point process models, which commonly assume independence between marks and locations—an assumption often violated in real-world applications such as forestry. To overcome this constraint, the authors propose a unified framework that, for the first time, enables comprehensive modeling, parameter estimation, simulation, and visualization of location-dependent marked point processes within the R programming environment. Grounded in spatial point process theory, the approach integrates statistical modeling with computational tools to support fitting to empirical data, model diagnostics, and generation of realistic spatial patterns. By relaxing the restrictive independence assumption, this work provides a practical and extensible analytical toolkit for researchers in ecology and related fields.
This work proposes a multiresolution cokriging model to address the challenge of sparse or missing observations in multivariate spatial prediction. By jointly modeling latent effects across multiple correlated spatial processes, the method enables information sharing and collaborative prediction. It integrates spatial basis functions with Gaussian Markov random field coefficients, explicitly capturing cross-process dependence without assuming global stationarity, while preserving the computational efficiency of fixed-rank kriging. Parameter estimation is carried out via an expectation–maximization algorithm. Simulation studies demonstrate that the approach effectively leverages densely observed variables to improve prediction accuracy in regions with sparse or no observations. The practical utility of the method is further validated through an application to PM10 concentration prediction in northern Italy.