S$^3$Geo: Structure-Semantic Synergistic Learning for Cross-View Geo-Localization

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of cross-view geo-localization to confusion between visually similar regions, a limitation arising from the decoupling of structural and semantic representations. To overcome this, we propose a structure-semantic collaborative learning framework. Specifically, a decoupled query pooling module is designed to extract local structural features, while optimal transport theory is leveraged to establish precise correspondences and enhance discrimination against hard negative samples. Furthermore, knowledge distillation from a frozen CLIP model is introduced to inject global semantic priors. Experimental results demonstrate that the proposed method surpasses existing state-of-the-art approaches on both the University-1652 and SUES-200 datasets without increasing inference complexity, effectively validating the superiority of joint structure-semantic modeling.
📝 Abstract
Cross-view geo-localization (CVGL) aims to estimate geographic locations by matching images captured from different viewpoints, such as drone and satellite views. Existing methods mainly rely on visual representations, but often fail to jointly model fine-grained structural correspondences and semantic priors, making them prone to confusion between visually similar but semantically different regions, and thus limiting robustness under large viewpoint variations. To address these challenges, we propose \textbf{S$^3$Geo}, a structure-semantic synergistic learning framework for cross-view matching. Specifically, we first introduce a Decoupled Query Pooling (DQP) module to extract a compact set of region-aware features from dense tokens, enabling explicit modeling of local structural patterns. We then design a query-level contrastive learning scheme with an optimal transport (OT)-based formulation to establish soft correspondences under cross-view spatial misalignment. Furthermore, we incorporate a Semantic Knowledge Distillation (SKD) strategy from a frozen CLIP teacher to transfer semantic priors and relational structures, thereby improving discrimination on hard negatives. By operating synergistically, the semantic priors provide robust contextual filtering, which guides the structural module to establish precise spatial alignments. Experiments on the University-1652 and SUES-200 datasets demonstrate that \textbf{S$^3$Geo} consistently outperforms state-of-the-art approaches without increasing inference complexity, validating the effectiveness of jointly modeling structural and semantic information for CVGL.
Problem

Research questions and friction points this paper is trying to address.

Cross-view geo-localization
Structural correspondences
Semantic priors
Viewpoint variations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-View Geo-Localization
Decoupled Query Pooling
Optimal Transport
Semantic Knowledge Distillation
Structure-Semantic Synergistic Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Ziqian Mo
The Hong Kong University of Science and Technology (Guangzhou)
H
Hill Zhang
Claremont McKenna College
Haosheng Tan
Haosheng Tan
University of Bristol
machine learningdeep learningmulti-modal learning
L
Ling Li
The Hong Kong University of Science and Technology (Guangzhou)
J
Jiaheng Wei
The Hong Kong University of Science and Technology (Guangzhou)