SAGE: A sampling-aware global evaluation benchmark for species distribution modeling
This study addresses the challenge that data biases frequently obscure performance disparities among rare species in multi-species distribution model evaluations. By integrating GBIF and sPlotOpen datasets, this work constructs a sampling-aware global evaluation benchmark and proposes a stratified assessment framework based on sampling effort and species prevalence. Spatial thinning and reweighting techniques are further implemented to correct for sampling bias. The results reveal that deep learning-based species distribution models (DeepSDMs) significantly outperform conventional single-species models for taxa with low-frequency occurrence records, underscoring the critical role of bias correction. Ultimately, this research provides a transparent and systematic evaluation foundation for ecologically credible modeling.