NavHarness: Adaptive Goals for Agentic Vision-Language Navigation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of misalignment between local actions and global routes, as well as excessive inference overhead caused by historical accumulation in long-horizon vision-language navigation. To this end, we propose a multi-agent collaborative framework comprising Goal, Verify, Memory, and Visuomotor agents. The framework introduces an adaptive goal-setting mechanism to dynamically align global planning with local execution, alongside a multimodal context compression strategy that effectively reduces computational redundancy. Experimental results demonstrate that the proposed method achieves superior performance on the R2R and RxR benchmarks, attaining an 83.3% success rate in real-world scenarios with a navigation error of only 1.51 meters. These findings indicate that our approach successfully balances navigational consistency with computational efficiency.
📝 Abstract
Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks. Moreover, the accumulated interaction history increases the input required for subsequent decisions, resulting in a significant inference overhead. To this end, we introduce \method, an Agentic VLN framework that includes a Goal Agent that sets adaptive goals for local actions, a Verify Agent that dynamically verifies whether a goal has been completed, a Memory Agent for multimodal context compression, and a Visuomotor Agent to execute adaptive goals. Specifically, the Goal Agent formulates adaptive goals based on the instruction, current observation, and execution history. Then the Visuomotor Agent executes navigation actions to achieve each goal, while the Verify Agent uses a goal-specific verification question to dynamically assess whether the observed outcomes satisfy the intended completion condition. Verified goal completion then marks a boundary for the Memory Agent to compress the corresponding multimodal interaction history while preserving information needed for subsequent navigation. We evaluate navigation on R2R-CE and RxR-CE, examine framework variants across three model backbones, and study context evolution during execution. For Real-World evaluation, \method achieves 83.3\% success and 1.51\,m navigation error across eight challenging routes evaluated three times each.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Navigation
Embodied Agents
Long-horizon Tasks
Inference Overhead
Multimodal Context Compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Navigation
Agentic Framework
Adaptive Goals
Context Compression
Multimodal Agents
🔎 Similar Papers