Improving Cross-view Object Geo-localization: A Dual Attention Approach with Cross-view Interaction and Multi-Scale Spatial Features

📅 2025-10-30
📈 Citations: 0
Influential: 0
📄 PDF

career value

194K/year
🤖 AI Summary
Existing cross-view geolocalization methods suffer from insufficient inter-view information interaction and coarse-grained spatial relationship modeling, making them vulnerable to edge noise. To address these issues, we propose a dual-attention mechanism: (1) a cross-view cross-attention module enabling bidirectional contextual modeling between aerial and ground views; and (2) a multi-head spatial attention module that fuses multi-scale convolutional features to strengthen implicit correspondence learning. Furthermore, we introduce the first fine-grained Ground-to-Drone (G2D) localization benchmark dataset. Extensive experiments on CVOGL and the proposed G2D dataset demonstrate that our method effectively suppresses irrelevant noise, enhances spatial relational representation, and achieves superior localization accuracy over state-of-the-art approaches.

Technology Category

Application Category

📝 Abstract
Cross-view object geo-localization has recently gained attention due to potential applications. Existing methods aim to capture spatial dependencies of query objects between different views through attention mechanisms to obtain spatial relationship feature maps, which are then used to predict object locations. Although promising, these approaches fail to effectively transfer information between views and do not further refine the spatial relationship feature maps. This results in the model erroneously focusing on irrelevant edge noise, thereby affecting localization performance. To address these limitations, we introduce a Cross-view and Cross-attention Module (CVCAM), which performs multiple iterations of interaction between the two views, enabling continuous exchange and learning of contextual information about the query object from both perspectives. This facilitates a deeper understanding of cross-view relationships while suppressing the edge noise unrelated to the query object. Furthermore, we integrate a Multi-head Spatial Attention Module (MHSAM), which employs convolutional kernels of various sizes to extract multi-scale spatial features from the feature maps containing implicit correspondences, further enhancing the feature representation of the query object. Additionally, given the scarcity of datasets for cross-view object geo-localization, we created a new dataset called G2D for the "Ground-to-Drone" localization task, enriching existing datasets and filling the gap in "Ground-to-Drone" localization task. Extensive experiments on the CVOGL and G2D datasets demonstrate that our proposed method achieves high localization accuracy, surpassing the current state-of-the-art.
Problem

Research questions and friction points this paper is trying to address.

Enhancing cross-view object geo-localization accuracy
Addressing ineffective information transfer between different views
Improving spatial feature representation with multi-scale attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual attention approach with cross-view interaction
Multi-scale spatial features extraction via attention
New dataset for ground-to-drone localization
🔎 Similar Papers
No similar papers found.
X
Xingtao Ling
College of Computer Science and Software Engineering, Shenzhen University, China
Y
Yingying Zhu
College of Computer Science and Software Engineering, Shenzhen University, China