Score
Transforming, aligning, and resampling geospatial datasets across coordinate systems, sensors, resolutions, and grids (including warping imagery into SAR slant-range or assembling multi-instrument products) to create consistent analysis-ready maps and pretraining datasets.
This work addresses three key challenges in multi-resolution SAR–optical remote sensing image registration: (1) mismatch of high-resolution structural details, (2) stereoscopic spatial misalignment, and (3) absence of standardized benchmark datasets. To this end, we present a systematic survey of existing methods, introduce MultiResSAR—the first publicly available multi-source, multi-resolution, multi-scene registration dataset comprising over 10,000 image pairs—and conduct a comprehensive evaluation of 16 state-of-the-art algorithms. Experimental results reveal a sharp degradation in registration accuracy with increasing resolution; all evaluated methods fail completely on sub-meter-resolution data. Among traditional methods, RIFT achieves the highest success rate (66.51%), while XoFTR leads among deep learning-based approaches (40.58%). Based on these findings, we propose three novel research directions: noise-robust feature suppression, 3D geometric fusion, and cross-view representation modeling—thereby establishing a foundational dataset, standardized evaluation protocol, and technical roadmap for high-resolution heterogeneous remote sensing image registration.
This work addresses the fragmentation and lack of standardized interfaces in the ecosystem of geospatial foundation model embeddings, which severely hinder model comparison and reproducibility. We formalize, for the first time, the Earth embedding product ecosystem and propose a three-tier taxonomy grounded in data, tools, and value dimensions. Through a systematic analysis of interoperability barriers, we extend TorchGeo to develop a unified API that treats Earth embeddings as standardized geospatial datasets, enabling plug-and-play integration of heterogeneous, multi-source embedding products. This framework effectively decouples downstream analytical tasks from embedding engineering, substantially lowering the barrier to entry and promoting reproducibility, transparency, and fair benchmarking in remote sensing workflows.
Remote sensing image fusion across heterogeneous satellite sensors (e.g., Landsat and Sentinel) remains challenging due to spectral response mismatches, temporal misalignment, and spatial resolution disparities. To address this, we propose the first end-to-end super-resolution framework designed explicitly for real-world multi-sensor data. Unlike conventional methods relying on synthetically degraded images, our approach directly aligns and reconstructs HLS30 imagery against high-resolution HLS10 reference data, explicitly modeling spectral–temporal inconsistencies. The framework integrates differentiable geometric registration with spectrum-aware reconstruction modules and is evaluated jointly via quantitative metrics (PSNR/SSIM) and qualitative assessment of structural fidelity. Experiments demonstrate that our method significantly improves spatial resolution consistency in HLS30 data, achieving an average PSNR gain of 2.1 dB while preserving spectral fidelity—establishing a new paradigm for practical super-resolution of heterogeneous remote sensing imagery.
This work addresses the prevalent lack of physical consistency in existing SAR-optical-text multimodal datasets, which are typically based on low-resolution intensity images and discard complex-valued measurements and native geometric structures. To overcome these limitations, we present the first large-scale, very-high-resolution (80 cm slant-range) SAR-optical-text triplet dataset, built upon open-source Umbra Spotlight data. Pixel-level geometric alignment is achieved through band-limited FFT resampling and local coordinate registration, preserving both complex SAR data and slant-range geometry. An automated pipeline generates hierarchical textual descriptions at SHORT, MID, and LONG levels. Encompassing 119,566 samples across 257 locations in 72 countries, the dataset supports cross-modal retrieval and conditional generation tasks, and includes standardized splits and baseline code.
To address the challenges of inconsistent feature representations across remote sensing sensors (e.g., Sentinel-2 and aerial imagery) and limited high-resolution annotated data—leading to poor cross-resolution generalization—this paper proposes X-STARS, a cross-sensor self-supervised training and alignment framework. Its core innovation is the first multi-sensor alignment dense loss, which achieves cross-resolution and cross-platform feature alignment via contrastive image-patch matching. X-STARS supports both from-scratch pretraining and continual pretraining paradigms. Evaluated on our newly constructed Cities-France multi-sensor dataset, X-STARS consistently outperforms state-of-the-art methods across seven downstream classification and segmentation tasks. Moreover, it achieves comparable performance using 30–50% fewer annotated samples, significantly reducing annotation burden while enhancing model transferability across heterogeneous remote sensing modalities.
This work addresses non-uniform, nonlinear spatial misalignment between post-disaster orthoimagery from small Unmanned Aerial Systems (sUAS) and prior building vector polygons—a critical yet previously unquantified challenge. Method: We conduct the first large-scale quantitative analysis across 51 orthoimages from nine disasters and 21,600 buildings, revealing an average translational error of 82 pixels and mean IoU of only 0.65; remarkably low angular and distance variances (0.4° and 0.45 px) confirm the absence of spatial consistency—invalidating the common assumption that linear transformations suffice, as often assumed in satellite remote sensing. We propose a rigorous validation framework integrating geometric metrics, IoU-based assessment, and GIS overlay comparison, and introduce the first publicly available benchmark dataset with manually refined ground-truth annotations. Contribution/Results: Our findings expose significant bias risks for downstream AI systems and establish a reproducible, paradigm-shifting foundation for sUAS georegistration research.
Remote sensing foundation models are hindered by small-scale, geographically narrow, and single-modality training datasets, limiting label-efficient large-scale pretraining. To address this, we introduce GeoEarth—the first global-scale, multimodal, spatiotemporally aligned Earth observation dataset—integrating eight modalities: optical, SAR, digital elevation, land cover, and others, spanning over 9 million globally distributed samples. GeoEarth is the first to systematically achieve co-registration, standardization, and spatiotemporal alignment of Analysis-Ready Data (ARD) across all eight modalities at planetary scale, thereby overcoming critical bottlenecks in modality diversity, geographic coverage, and data readiness. Extensive experiments demonstrate substantial performance gains on downstream tasks—including land-cover classification and change detection. The dataset is released with comprehensive metadata, detailed processing documentation, benchmark pretraining protocols, and a permissive open-source license.
This study addresses the lack of systematic best practices in large-scale Earth observation (EO) mapping, which often introduces errors during data preprocessing, model training, inference deployment, and validation, thereby compromising the reliability and scientific credibility of map products. To remedy this, we propose the first end-to-end best practice framework for EO mapping, encompassing the entire workflow from satellite data acquisition to operational map delivery. The framework integrates six core components: EO data infrastructure, preprocessing, machine learning dataset construction, uncertainty quantification, map production and dissemination, and independent validation. Emphasizing the interdependence of these stages, it embeds uncertainty quantification and independent validation as integral elements. By synergizing machine learning, distributed computing, and geospatial validation techniques, the framework establishes a reproducible and scalable mapping pipeline that substantially enhances the quality, consistency, and scientific rigor of EO-derived maps, supported by open-source resources to foster community adoption.
This work addresses the limitations of existing remote sensing datasets—such as single-resolution coverage, insufficient scale, and low alignment accuracy—that hinder the development of multi-scale, multi-modal foundation models. We present the first large-scale dataset comprising over 1.3 million pixel-level precisely aligned SAR and optical image pairs, spanning resolutions from 0.5 m to 10 m and covering 12 representative land-cover classes. A coarse-to-fine matching framework is introduced to effectively resolve challenges posed by multi-modal projection distortions and massive-scale registration, integrating multi-source data from Sentinel-1, PIESAT-1, Capella Space, and Google Earth. Comprehensive benchmarks across four vision tasks demonstrate significant performance gains, with state-of-the-art results in multi-modal matching, thereby filling a critical gap in high-precision, large-scale multi-modal remote sensing datasets.
High-resolution flood mapping is often hindered by cloud cover in optical imagery and speckle noise along with georegistration errors in synthetic aperture radar (SAR) data. To address these challenges, this work proposes a multimodal flood mapping framework that integrates Sentinel-1 and Sentinel-2 observations, introducing three key innovations: a translation-invariant loss function robust to registration offsets, a generative SAR despeckling model based on a conditional variational autoencoder (CVAE), and a weakly supervised label transfer strategy. Evaluated on a newly curated, high-quality multisensor flood dataset covering the conterminous United States, the proposed method achieves substantially improved accuracy under complex weather and urban conditions, attaining a multispectral area under the precision–recall curve (AUPRC) of 0.956 and demonstrating markedly superior SAR-based mapping performance compared to conventional filtering approaches.
This work addresses the degradation of localization accuracy in visual-inertial systems during long-term aerial deployments, which arises from calibration drift. To mitigate this issue, the authors propose a self-calibration method leveraging publicly available geospatial reference data—specifically, satellite orthoimagery and digital elevation models—as a global geometric prior. By integrating 2D–3D feature matching with joint visual-inertial optimization, the approach automatically refines both intrinsic and extrinsic camera parameters without requiring dedicated calibration maneuvers or manually placed ground control points. Evaluated over six large-scale aerial missions spanning two years, the method consistently outperforms established baselines including Kalibr, COLMAP, and VINS-Mono, significantly reducing reprojection and pose estimation errors while enhancing the long-term accuracy and robustness of the system in real-world operational scenarios.
Geospatial foundation models (GFMs) face two critical deployment bottlenecks: lack of automated data processing and model bloat after fine-tuning. This paper proposes an end-to-end geospatial machine learning framework integrating automatic multispectral image annotation, unified data pipeline orchestration, task-aware knowledge distillation, and lightweight model architecture—designed for open-source remote sensing data (e.g., Landsat, Sentinel-2). Our approach reduces model size by 8× and significantly cuts carbon footprint while preserving or improving accuracy: crop segmentation achieves 60.65% mIoU—12 percentage points above state-of-the-art—and matches or exceeds baseline performance in flood mapping and desert locust forecasting. The entire workflow—from raw imagery to web-map integration—is completed within 24 hours, substantially enhancing GFMs’ practicality and real-world deployability.