Score
Building models that jointly represent spatial and temporal structure (and their fusion with other modalities) to perform tasks like video-based clinical measurement, climate/cryosphere forecasting, or unified Earth-system representation across space and time.
Multimodal geospatial foundation models (GFMs) face core challenges in remote sensing—including modality heterogeneity, distribution shift, and semantic gap—hindering robust cross-domain generalization and interpretability. Method: This paper presents the first systematic, modality-driven survey of their technical evolution, proposing cross-modal interaction design principles grounded in remote sensing physics. Our approach integrates multi-resolution and multi-temporal modeling, contrastive learning, cross-modal attention, and task-adaptive fine-tuning to establish a unified alignment–transfer–generalization pipeline. Contribution/Results: We introduce the first taxonomy specifically for multimodal GFMs in remote sensing, release several novel benchmark datasets, and comprehensively evaluate dozens of models across ten downstream tasks—including land cover mapping, agricultural monitoring, and disaster response. Experiments demonstrate substantial improvements in cross-domain generalization and semantic interpretability, advancing the practical deployment of multimodal GFMs in geospatial AI.
Remote sensing foundation models hold significant promise for geospatial intelligent interpretation, yet their practical deployment remains hindered by challenges including difficulty in fusing heterogeneous multi-source data, scarcity of high-quality annotations, immature cross-modal alignment mechanisms, and prohibitive computational overhead. To address these issues, this work presents a systematic survey of vision- and multimodal-based remote sensing foundation models. It is the first to comprehensively analyze multimodal alignment, cross-modal transfer, and scalability challenges across optical, SAR, LiDAR, textual, and geospatial modalities. We propose a pragmatic model evolution roadmap and an open resource ecosystem, incorporating ViT architectures, CLIP-style contrastive learning, self-supervised pretraining, and remote sensing–specific data curation strategies. A unified evaluation framework is established, synthesizing over 120 models and datasets; all resources are open-sourced via GitHub, providing an authoritative benchmark and practical guidance for remote sensing large models.
Remote sensing time-series analysis faces challenges of fragmented multi-task modeling and difficulty in unifying spatiotemporal feature representation. To address this, we propose the first general-purpose generative framework supporting reconstruction, cloud removal, change detection, and forecasting. Built upon the flow-matching paradigm, our architecture integrates a diffusion-based Transformer with two novel components: an Adaptive Conditional Injector (ACor) and a Spatiotemporal-aware Modulator (STM), enabling joint modeling of multimodal conditional awareness and long-range spatiotemporal dependencies. Extensive experiments demonstrate significant superiority over state-of-the-art methods under challenging scenarios—including severe cloud contamination, missing modalities, and phenological forecasting. Furthermore, we release two high-quality multimodal remote sensing time-series datasets—TS-S12 and TS-S12CR—establishing new benchmarks and paving the way for unified time-series modeling in remote sensing.
This paper addresses the challenge of effectively fusing time-series and single-temporal remote sensing imagery. We propose a task-agnostic multimodal unified representation framework. Methodologically, we introduce two novel components: (1) temporal discretization and cross-modal token alignment, and (2) a hybrid quantization strategy combining deterministic and learnable quantization, coupled with a masked correlation learning objective. These enable semantic alignment of images and time series within a shared embedding space, supporting cross-modal generation—e.g., counterfactual reasoning and global temperature field reconstruction—as well as downstream task transfer. Evaluated on multiple remote sensing benchmarks, our pretrained model achieves an average R² improvement of 6 percentage points (+50% relative to baseline) and a 2-percentage-point reduction in RMSE (−12% relative to baseline), demonstrating substantial gains in generalization and robustness.
Existing Earth observation foundation models are constrained by fixed spatiotemporal scales, struggling to simultaneously achieve high spatial detail and high temporal fidelity. To address this, we propose a two-stage multimodal representation learning framework that jointly models Sentinel-1 (SAR) and Sentinel-2 (optical) data. While preserving the architectural independence of each modality’s encoder, our method leverages self-supervised learning and a cross-modal fusion network to construct a shared, consistent high-resolution feature space. The resulting embeddings achieve 10-meter spatial resolution, cloud-free reconstruction, and daily temporal continuity. Evaluated on global gross primary production (GPP) modeling, our approach significantly improves ecological interpretability and spatiotemporal consistency. This work marks the first successful synergistic representation of multi-source remote sensing data for fine-grained ecosystem dynamic modeling.
To address the challenges of modeling Earth observation (EO) multimodal data—particularly the difficulty in jointly capturing fine-grained spatial details and high-level semantics—this paper introduces the first generative multimodal foundation model for EO supporting arbitrary modality-to-arbitrary modality translation. Methodologically, we propose a novel dual-scale (token-level + pixel-level) early-fusion pretraining paradigm, jointly trained on nine global geospatial modalities; we further introduce “Thinking-in-Modality” (TiM), a mechanism enabling dynamic in-modal sample augmentation during inference and fine-tuning. Our contributions include: (1) open-sourcing both the model weights and a high-quality, large-scale EO multimodal dataset; and (2) achieving state-of-the-art performance across standard benchmarks (e.g., PANGAEA), unifying cross-modal generation, semantic understanding, and spatial reasoning within a single framework, while significantly improving zero-shot and few-shot generalization capabilities.
Remote sensing foundation models are hindered by small-scale, geographically narrow, and single-modality training datasets, limiting label-efficient large-scale pretraining. To address this, we introduce GeoEarth—the first global-scale, multimodal, spatiotemporally aligned Earth observation dataset—integrating eight modalities: optical, SAR, digital elevation, land cover, and others, spanning over 9 million globally distributed samples. GeoEarth is the first to systematically achieve co-registration, standardization, and spatiotemporal alignment of Analysis-Ready Data (ARD) across all eight modalities at planetary scale, thereby overcoming critical bottlenecks in modality diversity, geographic coverage, and data readiness. Extensive experiments demonstrate substantial performance gains on downstream tasks—including land-cover classification and change detection. The dataset is released with comprehensive metadata, detailed processing documentation, benchmark pretraining protocols, and a permissive open-source license.
This work addresses the challenges of modeling multimodal, multi-temporal Earth observation data with variable input lengths and the failure of existing foundation models in temporal tasks such as natural disaster risk prediction. To overcome these limitations, we propose TerraFlow, a novel approach that introduces a temporally oriented training objective to jointly model spatial, temporal, and modal dimensions. TerraFlow is the first method to enable unified representation learning for variable-length multimodal Earth observation sequences through a sequence-aware architecture that effectively fuses multimodal information and captures temporal dynamics. Evaluated on the GEO-Bench-2 benchmark, TerraFlow consistently outperforms current state-of-the-art models across all temporal tasks, achieving up to a 50% improvement in F1 score and a 24% reduction in Brier score, thereby resolving the catastrophic failures of prior methods in disaster risk mapping.
This work addresses the challenge of high-precision multimodal spatiotemporal modeling in planetary-scale dynamical systems—such as ecosystems—where large-scale annotated data are scarce. We propose DeepEarth, a self-supervised multimodal world model featuring the novel Earth4D encoder, which extends multi-resolution hash embeddings into the temporal dimension to enable global 4D modeling at sub-meter spatial and sub-second temporal resolution across century-scale timeframes. A learnable hash probe is introduced to enhance data efficiency. By integrating multimodal signals—including vision and language—through a masked reconstruction objective, DeepEarth achieves state-of-the-art performance on ecological forecasting benchmarks, substantially outperforming existing large multimodal foundation models. The code and models are publicly released.
This study addresses the pronounced spatiotemporal heterogeneity in the reliability of subseasonal-to-seasonal (S2S) temperature forecasts, which cannot be adequately captured by conventional approaches relying solely on lead time. The authors propose a dual-scale learning framework that disentangles calendar-aligned climatic background states from lead-time-matched recent weather evolution. By incorporating a spatially adaptive fusion mechanism and topology-aware constraints, the model jointly represents multiscale temporal components, spatial heterogeneity, and large-scale climate modes. The approach substantially enhances forecast stability at 30–90-day lead times, particularly over high-latitude regions and complex terrain during winter. Moreover, the learned fusion weights reveal a reorganization of predictability governed by seasonal and geographic factors, thereby reshaping the conceptual paradigm of S2S predictability.
This study addresses the challenge of continuous spatiotemporal prediction arising from spatial support mismatch between point observations and gridded data, as well as spatiotemporal misalignment of covariates. We propose a hierarchical Bayesian spatiotemporal fusion framework. Methodologically, we formulate a latent Gaussian field model grounded in the Matérn stochastic partial differential equation (SPDE) prior to jointly integrate heterogeneous multi-source data; we innovatively enable joint modeling of misaligned covariates and achieve efficient Bayesian inference via integrated nested Laplace approximation (INLA) coupled with the SPDE approach. Applied to daily soil moisture prediction across Scotland, the framework produces high-resolution continuous spatiotemporal maps, significantly outperforming single-source models in accuracy while delivering full uncertainty quantification. It establishes a scalable, interpretable statistical modeling paradigm for cross-scale environmental monitoring.
This work addresses the challenge of integrating heterogeneous, high-dimensional data inherent in interdisciplinary scientific problems, where existing AI models are often confined to single modalities and struggle to unify understanding and generation across diverse scientific sources. To this end, we propose FuXi-Uni, the first general-purpose framework that natively unifies multimodal scientific data understanding and generation within a shared architecture. By aligning scientific tokens with natural language and employing a dedicated scientific decoder, FuXi-Uni constructs a shared latent space that preserves both cross-disciplinary generality and domain-specific performance. The framework achieves state-of-the-art results in Earth system modeling, including 10-day global weather forecasting, tropical cyclone track and intensity prediction, and super-resolution downscaling, while also outperforming leading multimodal large language models on biomedical visual question answering benchmarks.