🤖 AI Summary
This study addresses the lack of large-scale, high-quality, multimodal benchmark datasets that has hindered systematic evaluation of the generalization capabilities of geospatial foundation models for flood segmentation. We introduce GEOID-Flood, the first dataset integrating dual-temporal Sentinel-1 (GRD/RTC), Sentinel-2 composite imagery, and digital elevation models (DEMs) from 219 flood events, providing over 14,000 precisely co-registered image patches with expert-validated labels. Leveraging this dataset, we evaluate prominent foundation models under single-image, multi-temporal, and multimodal settings, finding that fine-tuning with optical-SAR fusion most effectively captures instantaneous flood extents. Models trained on GEOID-Flood significantly outperform existing approaches on unseen flood events, demonstrating strong cross-regional and cross-sensor generalization.
📝 Abstract
Geospatial foundation models aim to learn representations that transfer across regions and sensors, yet evaluating them on specific tasks requires large, high-quality, multi-modal benchmarks that measure how well such models extract value from data. Concerning flood mapping, existing datasets rarely combine bi-temporal SAR and co-registered optical imagery at scale, leaving the value of foundation models for this downstream task largely untested. We introduce GEOID-Flood, a large-scale multi-modal flood segmentation benchmark, derived from Copernicus Emergency Management Service activations, spanning 219 events across 65 countries over ten years. The dataset provides more than 14,000 tiles with co-registered pre- and post-event Sentinel-1, in GRD and RTC format, pre-event Sentinel-2 composite, and DEM, including manually validated labels that separate background from permanent water and flooded water. Using this benchmark, we evaluate foundation models against conventional encoders across single-image, multi-temporal, and multi-modal protocols. We report three main findings: foundation models offer a consistent but modest advantage; optical-SAR fusion with finetuning best resolves transient flooding; and models trained on GEOID-Flood transfer to unseen events better than those trained on existing datasets. Dataset and code available at https://github.com/links-ads/geoid-flood.