Score
Designs and evaluates methods, models, and pipelines that align geospatial data and reconstructed spatial representations to semantic priors or manifolds so that geometry, features, and labels are consistent across modalities. This includes algorithms and transformations that constrain reconstructions, preserve spatial–semantic structure, and increase robustness of downstream interpretation and analysis.
This work addresses the limitations of existing Earth observation foundation models, which predominantly rely on raster data and overlook the structured geographic semantics embedded in open vector datasets such as OpenStreetMap, thereby hindering comprehensive understanding of human–environment systems. To overcome this, we propose the first unified spatial representation learning framework that deeply integrates remote sensing imagery and vector data within a shared embedding space, breaking away from conventional modality-isolated paradigms. By leveraging self-supervised learning and multimodal alignment—while explicitly modeling geometric, topological, and semantic relationships—our approach enables synergistic raster perception and vector-based reasoning. The method substantially enhances accuracy, semantic interpretability, and explainability on downstream tasks, laying a theoretical and methodological foundation for developing human-centered, semantically rich geospatial foundation models.
This work addresses the challenge of reliably inferring complete 3D geometry and semantics from in-vehicle images, which is hindered by occlusions, limited field of view, and insufficient constraints in unobserved regions. To this end, the authors propose GeoScene, a novel framework that, for the first time, incorporates structured geospatial priors—such as road and building layouts derived from satellite imagery and OpenStreetMap—as learnable soft constraints. The method adaptively fuses local visual observations with global structural knowledge through voxel-wise reliability weighting and introduces a prior-guided feature refinement network. Evaluated on SemanticKITTI and SSCBench-KITTI-360, GeoScene significantly advances 3D semantic scene completion performance, particularly excelling on large-scale static objects and geographically structured categories.
To address high geometric ambiguity and weak semantic guidance in indoor 3D reconstruction from sparse views, this paper proposes the first end-to-end geometric-semantic co-optimization framework. Methodologically, it integrates semantic priors from 2D foundation models and introduces depth-consistency constraints alongside multi-face normal regularization, enabling semantics to actively guide geometric optimization. By unifying differentiable rendering with multi-view geometry, the framework performs semantic-driven optimization of 3D Gaussian splatting. Evaluated on standard benchmarks including ScanNet, our method achieves state-of-the-art performance in both novel-view synthesis and geometric reconstruction. It significantly improves model completeness under sparse input—reducing Chamfer distance by +12.3%—and enhances geometric fidelity—increasing PSNR by +8.7%. These results empirically validate the substantial benefit of semantic priors in strengthening geometric reconstruction accuracy and robustness.
Existing general-purpose multimodal embedding models struggle to effectively handle heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal dynamics. To address this limitation, this work proposes Geo-Embed, a unified embedding framework that leverages an instruction-tuned vision-language backbone to integrate street-view imagery, remote sensing data, textual descriptions, region masks, and temporal information, enabling versatile multimodal query-target matching. Concurrently, we introduce GeoMEB, the first large-scale multimodal embedding benchmark tailored for urban understanding, encompassing 45 diverse tasks. Experimental results demonstrate that Geo-Embed outperforms the strongest baseline by 15.3% overall on GeoMEB, achieving significant improvements across retrieval, visual question answering, change detection, classification, and visual grounding tasks.
Geographic spatial object data exhibit strong heterogeneity and suffer from severe label scarcity, hindering supervised learning. Method: This paper presents a systematic survey of self-supervised representation learning methods for three fundamental vector geometric primitives—points, lines, and polygons—unifying both predictive and contrastive paradigms. Contribution/Results: It introduces the first taxonomy organized by geometric type, encompassing over 100 state-of-the-art works; identifies seven key adaptation strategies and characterizes their effectiveness boundaries under multi-source data fusion and sparse-labeling conditions; and proposes an evolutionary pathway from task-specific models toward geographic foundation models, highlighting core challenges including cross-modal alignment and spatiotemporal consistency. The work establishes a theoretical framework and practical guidelines for GeoAI self-supervised modeling, enabling diverse downstream geospatial applications.
This work addresses the limitations of multimodal large language models (MLLMs) in spatial reasoning tasks, which stem from their reliance on static, single-layer geometric feature extraction and hinder diverse comprehension capabilities. To overcome this, the authors propose GeoAlign, a novel framework that introduces, for the first time, a dynamic multi-layer geometric feature alignment mechanism. GeoAlign constructs a hierarchical geometric feature bank and employs the MLLM’s original visual tokens as content-aware queries, combined with inter-layer sparse routing, to adaptively select and fuse multi-scale geometric information across image regions. This approach transcends the constraints of single-layer features by enabling task-driven geometric-semantic alignment. Experimental results demonstrate that GeoAlign achieves state-of-the-art performance on benchmarks including VSI-Bench, ScanQA, and SQA3D, with its 4B-parameter variant even outperforming larger existing MLLMs.
This work addresses the physical inconsistency in conventional multi-view satellite image evaluation, which relies on unconstrained 2D matching and ignores the epipolar geometry implicitly encoded in Rational Polynomial Coefficients (RPCs). The paper proposes the first geometry-aware evaluation protocol tailored to the RPC framework: it constructs a geometrically constrained search manifold via 3D projection and employs dense matching as a proxy task to assess the local uniqueness of features within a physically plausible space. By integrating geometric constraints into foundational model evaluation for the first time, this approach reveals a decoupling between semantic consistency and geometric localization capability, and establishes a reproducible, geometry-faithful benchmark for satellite imagery. Experiments demonstrate that, under RPC-consistent evaluation, generic 2D backbone networks outperform specialized 3D-aware models, underscoring the fundamental importance of geometric constraints in task formulation.
This work addresses the challenge of unified modeling for multi-class heterogeneous geographic entities in remote sensing vector mapping, where existing methods struggle to adequately represent topological relationships and instance boundaries. The authors propose reframing vector map construction as a structured text generation task by designing a GeoJSON-like hierarchical vector language that jointly encodes geometric, semantic, and topological information. A progressive vision-to-language mapping framework is introduced, optimized via reinforcement learning to ensure syntactic validity, content fidelity, and map executability of the generated output, thereby enabling cross-category unified modeling. Experiments on the newly curated VecMap-Bench dataset—comprising 54K images and 800K instances—demonstrate that the proposed approach significantly outperforms state-of-the-art methods in single- and multi-class mapping, cross-dataset transfer, and open-vocabulary generalization.
This study addresses the unclear geometric structure embedded in Earth observation foundation models—such as AlphaEarth—and its impact on environmental reasoning. We reveal for the first time that their embedding manifolds exhibit non-Euclidean characteristics and quantify the relationships among effective dimensionality, local directional variation, and retrieval coherence. Building on these insights, we propose a retrieval-augmented reasoning framework that integrates manifold geometric analysis with multi-tool agents, enabling query decomposition and chain-of-thought reasoning through tangent space alignment and efficient FAISS-based retrieval. Experiments demonstrate that our approach significantly improves response quality, achieving an average score of 3.79 compared to 3.03 for baseline methods, and reaching 4.28 on multi-step comparison tasks. Moreover, higher-capability models benefit more from geometry-aware representations.
This work addresses the challenge of maintaining semantic, structural, and geometric consistency under extreme viewpoint variations in cross-view localization. The authors propose CROSS, a novel framework that reframes cross-view localization as a joint learning task beyond pose estimation, integrating 3D grounding alignment, structure-aware matching, and relative hypothesis ranking to cohesively model semantic, structural, and geometric consistency. Notably, CROSS introduces structure learning as an intrinsic constraint to preserve semantic integrity—avoiding the pitfalls of point-wise matching—and leverages the powerful 2D representation capabilities of vision foundation models to enhance geometric reasoning. Evaluated on KITTI and VIGOR benchmarks, the method achieves state-of-the-art performance, significantly improving semantic stability, structural reliability, and geometric transferability across drastically different viewpoints.