π€ AI Summary
Existing generative models for Earth observation rely on natural image priors, which struggle to satisfy geospatial constraints and suffer from poor scalability and viewpoint bias. This work proposes GeoCore-9B, the first 9-billion-parameter generative foundation model trained entirely from scratch on pure Earth observation data. Built upon a Flow Matching diffusion Transformer architecture, it enables joint conditional generation guided by text prompts and continuous geospatial metadataβsuch as latitude, longitude, and ground sampling distance. A novel geospatial semantic alignment loss is introduced, leveraging a frozen teacher network to distill surface structure priors and steer the diffusion trajectory without incurring additional inference overhead. Pretrained on the Git-10M global dataset, GeoCore-9B achieves new state-of-the-art performance in both visual fidelity and geospatial structural accuracy, and demonstrates strong capabilities in challenging tasks including cloud removal and cross-modal translation from SAR to optical imagery.
π Abstract
Existing generative models for earth observation (EO) predominantly rely on fine-tuning natural image priors, which limits their scalability and introduces perspective biases that conflict with geospatial constraints. To address this, we introduce GeoCore-9B, a 9-billion-parameter generative foundation model, which is the first of its scale to be trained from scratch exclusively on EO data. Unlike previous EO foundation models, GeoCore-9B is built upon a Flow Matching-based Diffusion Transformer (DiT) and natively conditions generation on text descriptions and continuous geospatial metadata, including ground sample distances, latitudes, and longitudes. To overcome the convergence and spatial disorientation challenges of training at this scale, we propose a Geospatial Semantic Alignment loss. This objective distills structural Earth surface priors (e.g., terrain and urban areas) from a frozen specialist teacher network, constraining the diffusion latent trajectory during training without adding inference overhead. Pre-trained on the global-scale Git-10M dataset, GeoCore-9B demonstrates strong downstream versatility. Beyond standard proxy generative tasks, we show that GeoCore-9B can be effectively adapted for practical EO applications, including highly challenging tasks such as cloud removal and SAR-to-optical cross-modal translation. Extensive evaluations confirm that GeoCore-9B establishes new state-of-the-art performance in both visual fidelity and geographic structural accuracy.