🤖 AI Summary
This study addresses the challenges of poor directional observability, difficult deviation recovery, and unstable termination in round-trip vision-language navigation. To this end, we propose a reliability-aware sparse path memory mechanism built upon the NaVILA model and the Unitree Go2 simulation platform. The method stores outbound trajectories via geometric anchors to facilitate reverse retrieval, while incorporating structured prompting, action-level arbitration, and terminal verification to effectively distinguish information quality from online reliability, thereby enabling continuous round-trip navigation. Experimental results demonstrate that the proposed system achieves an online success rate of 55.1%, significantly outperforming the purely language-based baseline (22.0%) and approaching the oracle upper bound with perfect information (86.0%).
📝 Abstract
Vision-language navigation (VLN) is typically evaluated as a one-way task, although deployed robots may need to return after reaching a goal. We study continuous round-trip VLN and diagnose failures in directional observability, deviation recovery, and termination stability. We propose a reliability-aware sparse route memory that records the executed Outbound trajectory as ordered geometric anchors and queries them in reverse through a structured hint, action-level arbitration, and terminal verification. On 50 reverse-paired episodes using NaVILA and a simulated Unitree Go2, language-only Return succeeds in 22.0% of episodes, while our online system reaches 55.1%. With exact route information, the same interfaces achieve 86.0%, showing that effective Return requires both accurate information and consistent action on that information. The remaining online gap arises mainly from geometric evidence that is too unreliable to authorise intervention. These results distinguish information quality, behavioural consistency, and online reliability as separate limits in long-horizon navigation.