🤖 AI Summary
This study addresses the challenge that data biases frequently obscure performance disparities among rare species in multi-species distribution model evaluations. By integrating GBIF and sPlotOpen datasets, this work constructs a sampling-aware global evaluation benchmark and proposes a stratified assessment framework based on sampling effort and species prevalence. Spatial thinning and reweighting techniques are further implemented to correct for sampling bias. The results reveal that deep learning-based species distribution models (DeepSDMs) significantly outperform conventional single-species models for taxa with low-frequency occurrence records, underscoring the critical role of bias correction. Ultimately, this research provides a transparent and systematic evaluation foundation for ecologically credible modeling.
📝 Abstract
Knowing where species occur is fundamental for biodiversity research and conservation. Species distribution models (SDMs) link species observations to environmental conditions to estimate their spatial distribution. However, accuracy varies with the underlying data and models, making it essential to know for which species models can be trusted. Deep-learning-based SDMs ("DeepSDMs") now jointly model thousands of species, drawing on hundreds of millions of community-science records. At this scale, averaging performance hides substantial species-level variability, particularly for rare species, often of greatest conservation concern. Records are also strongly biased, making occurrence counts misleading. Accounting for these factors is essential for a reliable and informative evaluation of multi-species SDMs. Here, we introduce a Sampling-Aware Global Evaluation (SAGE) benchmark, combining GBIF records for training with sPlotOpen vegetation plots for presence-absence evaluation across 5771 plant species. We propose an evaluation framework that groups species based on two properties, sampling effort and relative prevalence, which describe how densely a species' range is sampled and how frequently the species is recorded. Evaluating single-species SDMs and multi-species DeepSDMs, we find that Random Forests and DeepSDMs perform best overall, but neither dominates: DeepSDMs outperform single-species SDMs for infrequently recorded species while offering no consistent advantage for well-sampled ones. Crucially, this advantage emerges only when established bias-correction practices, such as spatial thinning and reweighting, are carried over to the deep-learning setting. SAGE helps identify the species and data conditions for which a given approach is beneficial, thereby supporting the development of more transparent and ecologically credible SDMs. Data and code: https://earens.github.io/sage/