🤖 AI Summary
Species distribution models (SDMs) suffer from spatial autocorrelation (SAC), leading to overly optimistic performance estimates under standard random cross-validation (CV) in spatiotemporal extrapolation. To address this, we systematically evaluate multiple CV strategies—spatial blocking and environmental clustering (ENV)—combined with distinct training paradigms (LAST FOLD vs. RETRAIN), and propose a novel spatiotemporal CV framework explicitly aligned with the intrinsic spatiotemporal structure of SDM data. We demonstrate for the first time that optimizing spatial blocking distance and applying ENV clustering significantly mitigate SAC-induced bias, while the LAST FOLD paradigm prevents SAC recurrence in validation sets, substantially reducing estimation error compared to conventional approaches. Empirical results show SP-422 and ENV achieve Spearman and Pearson correlations of 0.485 and 0.548, respectively, confirming that CV strategy must reflect the underlying spatiotemporal dependence structure. This work establishes a methodological benchmark for robust SDM evaluation.
📝 Abstract
Species Distribution Models (SDMs) often suffer from spatial autocorrelation (SAC), leading to biased performance estimates. We tested cross-validation (CV) strategies - random splits, spatial blocking with varied distances, environmental (ENV) clustering, and a novel spatio-temporal method - under two proposed training schemes: LAST FOLD, widely used in spatial CV at the cost of data loss, and RETRAIN, which maximizes data usage but risks reintroducing SAC. LAST FOLD consistently yielded lower errors and stronger correlations. Spatial blocking at an optimal distance (SP 422) and ENV performed best, achieving Spearman and Pearson correlations of 0.485 and 0.548, respectively, although ENV may be unsuitable for long-term forecasts involving major environmental shifts. A spatio-temporal approach yielded modest benefits in our moderately variable dataset, but may excel with stronger temporal changes. These findings highlight the need to align CV approaches with the spatial and temporal structure of SDM data, ensuring rigorous validation and reliable predictive outcomes.