Score
Designs and implements models and workflows to predict full probability distributions (or distribution-valued observations) at unsampled spatial locations, building methods such as distributional kriging and other spatial distribution prediction techniques. Builds and evaluates spatial-data models that explicitly model spatial dependence and covariate effects to produce distributional estimates and derived functionals (e.g., means, exceedance probabilities).
Existing surveys predominantly focus on single domains or model types, lacking interdisciplinary and systematic syntheses of spatiotemporal statistical models. Method: Guided by the PRISMA framework, we systematically identified and structurally analyzed 83 high-quality publications from 2021–2025 across epidemiology, ecology, public health, economics, and criminology. Contribution/Results: We propose a novel, unified classification framework for spatiotemporal models, revealing the ubiquity of hierarchical structures and additive spatiotemporal dependence. The analysis uncovers domain-specific modeling preferences and pronounced imbalances in research distribution. Critically, we find widespread deficiencies in reproducibility and methodological transparency—impeding cross-domain knowledge transfer. Our framework provides both a theoretical foundation and practical pathway for interdisciplinary model comparison, methodological borrowing, and principled model refinement.
Modeling large-scale, non-Gaussian, noisy, and incomplete spatial data remains challenging due to limitations of conventional copula models in capturing complex spatial heterogeneity and non-Gaussian dependence structures. Method: We propose a Bayesian hierarchical model integrating vine copulas with spatial random effects. Crucially, we design a novel vine copula structure explicitly embedding spatial dependence, enabling low-rank latent process representations and computationally efficient Bayesian inference. Contribution/Results: Our method overcomes key bottlenecks in traditional copula-based spatial modeling. In both parameter estimation and spatial prediction tasks, it significantly outperforms benchmarks—including fixed-rank kriging (FRK)—in accuracy, convergence speed, and robustness. Applied to atmospheric methane concentration mapping over Australia’s Bowen Basin using Sentinel-5P satellite remote sensing data, the approach delivers superior predictive performance under missingness and noise. This work establishes a new paradigm for non-Gaussian spatial statistical modeling.
This study addresses the technical gap between domain experts and non-experts in geostatistics by proposing and implementing the first open-source, integrated spatial interpolation library. To bridge this gap, we develop a unified framework that synergistically combines classical geostatistical methods (e.g., ordinary and universal kriging) with modern machine learning techniques. Our approach employs ensemble learning to automatically select and optimize interpolation strategies, and integrates stochastic simulation for posterior distribution inference—enabling both point estimation and rigorous uncertainty quantification. Built on a hybrid Python/C++ architecture, the library balances usability, computational efficiency, and extensibility. Empirical evaluation across diverse real-world spatial datasets demonstrates accuracy competitive with or superior to conventional methods. Moreover, the library provides standardized APIs, comprehensive uncertainty assessment, and built-in visualization tools—collectively lowering the barrier to entry for spatial modeling.
The long-standing conflation of Kriging and Gaussian Process Regression (GPR) lacks rigorous mathematical correspondence, leading to conceptual ambiguity and misinterpretation of their equivalence. Method: Starting from first principles, we systematically derive the precise GPR formulations corresponding to Simple, Ordinary, and Universal Kriging—explicitly mapping their underlying assumptions, estimation criteria (least squares, best linear unbiased estimation, maximum likelihood estimation), and predictive frameworks. Contribution/Results: We establish for the first time the exact equivalence conditions between each Kriging variant and specific GPR configurations, clarifying fundamental similarities and differences in covariance modeling, mean structure specification, and statistical inference logic. Our analysis refutes the oversimplified claim of “full equivalence” and provides a unified theoretical framework bridging geostatistics and machine learning. This enables principled cross-disciplinary understanding, methodological integration, and informed model selection across spatial statistics and probabilistic machine learning.
To address computational bottlenecks in latent spatial field inference and pointwise prediction in geostatistical models, this paper proposes Bayesian Predictive Stacking (BPS). BPS analytically aggregates posterior distributions of regression coefficients and spatial processes across multiple hyperparameter configurations, circumventing iterative algorithms such as MCMC and enabling fully parallelized inference. Its key innovation lies in a unified stacking framework that jointly combines predictive means and posterior densities, underpinned by infill asymptotic theory that guarantees statistical consistency. Experiments demonstrate that BPS achieves predictive accuracy comparable to full-sample Bayesian inference while drastically reducing computational time. The method is thus highly efficient, robust to hyperparameter specification, and scalable to large spatial datasets.
Existing spatial data models struggle to simultaneously capture complex extremal features such as tail dependence, asymptotic independence, and tail asymmetry. This work proposes a novel approach by introducing multivariate Pareto mixture distributions into a spatial copula framework, yielding a flexible model capable of jointly modeling both bulk and tail behaviors. The resulting formulation provides a unified treatment of the three aforementioned extremal dependence structures while preserving permutation asymmetry. Based on copula theory and maximum likelihood estimation, the method is validated through finite-sample simulations that demonstrate its computational feasibility and favorable parameter estimation performance. Empirical analysis of temperature data successfully reveals intricate tail structures, underscoring the model’s theoretical rigor and practical utility.
Traditional geostatistical methods rely on second-order moments and Gaussian assumptions, which are inadequate for capturing non-Gaussian spatial dependence. This work addresses this limitation by leveraging Sklar’s theorem to introduce a copula-based framework that decouples marginal distributions from the spatial dependence structure. The authors systematically develop spatial copula models applicable at both fixed point sets and process levels, emphasizing Kolmogorov consistency to clarify distinctions between these two modeling paradigms. The framework is further extended to spatio-temporal settings and supports flexible marginal specifications. By integrating spatial statistics, copula theory, and stochastic processes, this study establishes a unified approach for modeling non-Gaussian spatial dependence, elucidating the relationships, strengths, and limitations of existing methodologies, thereby advancing both theoretical understanding and practical applications in the field.
This study addresses the lack of systematic comparisons among effective approaches for handling spatial autocorrelation in random forests. Integrating machine learning with geostatistical principles, it presents the first unified evaluation of four strategies—Gaussian process augmentation, observation-driven correlation structures, spatial basis functions, and local geographical fitting—to enhance the spatial predictive performance of random forests. Through both simulation experiments and an empirical analysis of air pollution in Blantyre, Malawi, the research demonstrates that while no single method consistently outperforms others across all scenarios, spatial basis functions exhibit robust and superior performance throughout. These findings offer practical and reliable guidance for spatial modeling of environmental processes using random forests.
This study addresses the prohibitive computational cost of traditional kriging when applied to large datasets and dense prediction grids, which hinders efficient uncertainty quantification. The authors propose a fast kriging method based on a local L-order neighborhood sparse approximation. By leveraging the covariance structure of stationary Gaussian processes and properties of the Matérn covariance function, the approach approximates off-grid covariance vectors as sparse linear combinations of nearby grid points. This yields substantial reductions in computational complexity while maintaining theoretically controlled approximation error. The method integrates seamlessly into conditional simulation frameworks, enabling efficient uncertainty propagation. In experiments with North American summer precipitation data, it achieves over 150× speedup compared to exact kriging, with prediction errors as low as approximately 10⁻⁵ inches—visually indistinguishable from exact results—and outperforms existing methods such as Vecchia approximation and LatticeKrig.
This study addresses the risk of inferential bias and underestimated uncertainty in traditional spatial statistics due to misspecification of covariance functions. To overcome this limitation, the authors propose a semiparametric Bayesian approach based on Bernstein polynomials that flexibly models spatial covariance structures without requiring a prespecified parametric form, making it suitable for georeferenced data with latent spatial effects, such as those arising in spatial generalized linear mixed models. Efficient posterior inference is achieved via Markov chain Monte Carlo (MCMC), and the framework naturally supports Bayesian kriging prediction. Simulation studies demonstrate that the method accurately recovers true covariance structures with low bias and high precision. Applied to North American robin abundance data, it successfully captures overdispersion, estimates a spatial range of 381.48 kilometers, and exhibits strong out-of-sample predictive performance.
This study addresses the long-standing question of whether multivariate Kriging outperforms single-output modeling under heterotopic observations. The work establishes, for the first time, a theoretical link between output-specific design geometry and the estimability of cross-output dependencies, yielding a model-free diagnostic criterion. It derives an exact expression for predictive gain along with its geometric bounds, leading to a practical first-order net benefit criterion. Theoretical analysis reveals the statistical nonequivalence between collocated and heterotopic designs. Extensive experiments—conducted using separable multi-output Gaussian processes, radial basis functions, and linear models of coregionalization—validate the proposed criterion on synthetic datasets, an M/M/1 queueing system, and a multi-pollutant monitoring network, demonstrating both its effectiveness and practical utility.