🤖 AI Summary
Standard leave-one-out cross-validation (LOO-CV) fails for long-range prediction evaluation in latent Gaussian models—especially those with structured random effects—because it implicitly assumes independent observations, violating the spatial, spatiotemporal, or network dependencies inherent in real-world extrapolation settings. To address this, we propose Group-Out Cross-Validation (Group-Out CV): a dependency-aware CV scheme that automatically identifies and excludes groups of statistically dependent observations, thereby emulating realistic out-of-sample prediction scenarios. We further introduce a Bayesian joint posterior correction method that adjusts predictive distributions without refitting the model. This constitutes the first CV framework explicitly designed for dependent data, featuring automatic grouping and no re-estimation. Implemented in the open-source R-INLA package, our approach demonstrably enhances the robustness and reliability of predictive assessment compared to LOO-CV, particularly in models with complex dependence structures.
📝 Abstract
Evaluating the predictive performance of a statistical model is commonly done using cross-validation. Although the leave-one-out method is frequently employed, its application is justified primarily for independent and identically distributed observations. However, this method tends to mimic interpolation rather than prediction when dealing with dependent observations. This paper proposes a modified cross-validation for dependent observations. This is achieved by excluding an automatically determined set of observations from the training set to mimic a more reasonable prediction scenario. Also, within the framework of latent Gaussian models, we illustrate a method to adjust the joint posterior for this modified cross-validation to avoid model refitting. This new approach is accessible in the R-INLA package (www.r-inla.org).