AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of benchmarks, inconsistent cross-platform spatial context, and low single-platform planning reliability in air-ground collaborative vision-language navigation. To tackle these challenges, we propose the AG-CoNAV framework and construct AirGroundVLN, a large-scale benchmark comprising tens of thousands of scenarios. Methodologically, we design a Spatiotemporal Anchored Collaborative Memory (SACM) to maintain cross-view consistency and introduce an Aerial-Guided Region-to-Local hierarchical Planning (AGRLP) strategy that integrates broad aerial guidance with fine-grained ground-level localization. Extensive experiments validate the effectiveness of the proposed framework, demonstrating significant improvements in navigation success rates. Furthermore, this work establishes a comprehensive benchmark for advancing research in air-ground collaborative navigation.
📝 Abstract
Goal-oriented Vision-and-Language Navigation (VLN) requires agents to locate and reach targets described in natural language without prescribed routes. Air--ground collaboration is valuable for tasks requiring both wide-area search and fine-grained localization. However, systematic study of goal-oriented air--ground collaborative VLN remains limited by the lack of large-scale, diverse benchmarks and two core challenges: 1) substantial differences between aerial and ground views, together with useful observations becoming unavailable as navigation proceeds, make it difficult to maintain spatially consistent context across platforms and over time; and 2) asymmetric spatial observability makes ground perception locally detailed but spatially limited and aerial perception broad but locally coarse, limiting the reliability of single-platform planning. To address these limitations, we introduce AirGroundVLN, a benchmark containing 10,281 navigation episodes and 955 target instances across 19 Unreal Engine environments, with seen/unseen splits and an aerial-visibility protocol for systematic evaluation. Alongside the benchmark, we propose AG-CoNAV, a trainable reference framework comprising two key components: Spatiotemporally Anchored Collaborative Memory (SACM) and Aerial-Guided Regional-to-Local Planning (AGRLP). SACM maintains and retrieves spatially consistent historical context across aerial and ground observations. Meanwhile, AGRLP combines regional aerial guidance with fine-grained ground navigation. Extensive experiments demonstrate the effectiveness of AG-CoNAV and establish AirGroundVLN as a comprehensive benchmark for future exploration.
Problem

Research questions and friction points this paper is trying to address.

Vision-and-Language Navigation
Air-Ground Collaboration
Goal-Oriented Navigation
Benchmark
Spatial Observability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Air-Ground Collaborative VLN
Spatiotemporally Anchored Collaborative Memory
Aerial-Guided Regional-to-Local Planning
Goal-Oriented Navigation
Benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhenxuan Zeng
School of Computer Science, Northwestern Polytechnical University, China.
Q
Qingle Wu
School of Computer Science, Northwestern Polytechnical University, China.
W
Wei Suo
School of Computer Science, Northwestern Polytechnical University, China.
M
Maojia Wu
School of Computer Science, Northwestern Polytechnical University, China.
B
Bairong Zhang
School of Computer Science, Northwestern Polytechnical University, China.
H
Hangzheng Yu
School of Computer Science, Northwestern Polytechnical University, China.
Peng Wang
Peng Wang
School of Computer Science, Northwestern Polytechnical University, China
Computer VisionMachine LearningArtificial Intelligence