Towards Active Cross-View Object Geo-Localization

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出ActiveGeo方法,通过主动选择视角和决定停止时机来优化跨视角目标地理定位,使用ActiveMoPT框架实现,并在实验中展示了优越性能。
📝 Abstract
Cross-view object geo-localization (CVOGL) typically assumes a fixed query image, overlooking the ability of mobile agents to actively acquire more informative observations. To address this limitation, we introduce Active Cross-View Object Geo-Localization (ActiveGeo), where an agent sequentially selects new viewpoints and determines when to stop, aiming to improve localization with minimal observations. We further propose ActiveMoPT, an ActiveGeo framework with three-stage training. First, Multi-View Prompt-Preserving Adaptation enables the model to aggregate multiple query views while reusing the initial prompt. Second, Trajectory-Guided Policy Initialization uses supervised agent trajectories to learn viewpoint selection and initial stopping behavior. Third, Cost-Aware Policy Refinement employs GRPO with a gain-cost reward to jointly optimize localization accuracy and observation efficiency. We also construct ActiveGeo-858, a zero-shot test set containing 858 scenes and 1,716 target annotations. Experiments show that ActiveMoPT achieves state-of-the-art performance on MoP-UAV using only 1.45 query views on average, and substantially outperforms previous CVOGL approaches under zero-shot evaluation on ActiveGeo-858.
Problem

Research questions and friction points this paper is trying to address.

Cross-View Object Geo-Localization
Active Observation
Viewpoint Selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

ActiveGeo
Multi-View Prompt-Preserving Adaptation
Trajectory-Guided Policy Initialization
Cost-Aware Policy Refinement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shunyu Yao
College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310027, China
X
Xiaohan Zhang
College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310027, China
Zhuoran Yang
Zhuoran Yang
Yale University
machine learningoptimizationreinforcement learningstatistics
H
Haoqi Lai
College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310027, China
Q
Qi Ming
College of Computer Science, Beijing University of Technology, Beijing 100124, China
X
Xiaoxi Hu
State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University, Beijing 100084, China
H
Hui-Liang Shen
College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310027, China
Si-Yuan Cao
Si-Yuan Cao
Zhejiang University
image alignmenthomography estimationimage fusionplace recognition