Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing general-purpose multimodal embedding models struggle to effectively handle heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal dynamics. To address this limitation, this work proposes Geo-Embed, a unified embedding framework that leverages an instruction-tuned vision-language backbone to integrate street-view imagery, remote sensing data, textual descriptions, region masks, and temporal information, enabling versatile multimodal query-target matching. Concurrently, we introduce GeoMEB, the first large-scale multimodal embedding benchmark tailored for urban understanding, encompassing 45 diverse tasks. Experimental results demonstrate that Geo-Embed outperforms the strongest baseline by 15.3% overall on GeoMEB, achieving significant improvements across retrieval, visual question answering, change detection, classification, and visual grounding tasks.
📝 Abstract
Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal changes. To address this gap, we make three key contributions. First, we introduce GeoMEB, a large-scale multimodal embedding benchmark that standardizes 45 urban evaluation tasks across retrieval, visual question answering, change detection, classification, and visual grounding, together with training collections comprising 1.32M examples and 286K evaluation queries. Second, we present Geo-Embed, a unified embedding model that adapts a shared vision-language backbone to instruction-conditioned query-target matching over heterogeneous geospatial inputs, including single images, multiple images, text, regions, and masks. On GeoMEB, Geo-Embed achieves the strongest overall performance among representative multimodal embedders, with a 15.3% relative improvement over the strongest baseline. These results motivate future geospatial embedders that organize training and evaluation around explicit query-target relations, including semantic, cross-view, region-level, and temporal correspondence.
Problem

Research questions and friction points this paper is trying to address.

multimodal embeddings
geospatial understanding
urban applications
heterogeneous data
temporal change
Innovation

Methods, ideas, or system contributions that make the work stand out.

unified multimodal embeddings
geospatial understanding
instruction-conditioned matching
heterogeneous data fusion
urban benchmarking