HAM-VLN: Harnessing Hierarchical Agentic Memory for Zero-Shot Vision-and-Language Navigation

πŸ“… 2026-07-31
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the memory and reasoning bottlenecks in zero-shot vision-and-language navigation caused by long-horizon image streams or dense maps. To overcome these challenges, the authors propose a decision-coupled, self-generated memory mechanism that constructs a depth-aware world graph and simultaneously records semantic and reflective information during action selection. They introduce a hierarchical memory architecture that requires no additional large-model invocations, integrating relevance, recency, and saliency scoring with topological expansion to enable efficient management of historical context. The method achieves success rates of 61.0%, 52.7%, and 79.7% on VLN-CE R2R, RxR, and HM3D-v2 ObjectNav benchmarks, respectively, while compressing context length by over 65%.
πŸ“ Abstract
Vision-and-language navigation (VLN) enables robots to follow instructions in previously unseen environments. Recently, a training-free paradigm has emerged: the robot queries a multimodal LLM to understand its observations and plan the next action. However, long-horizon navigation based on either image streams or dense map inevitably introduces a growing memory and reasoning bottleneck. We present HAM-VLN, a decision-coupled, agent-authored memory that equips the robot with a persistent, depth-grounded world graph. In the same model call used to select the next action, HAM-VLN also records semantic and reflective information---including room type, objects, navigation progress, and failure notes. Recent waypoints remain verbatim within a bounded window, while older history re-enters the context only through retrieval scored by relevance, recency, and salience, together with one-hop topological expansion. This design requires no additional LLM calls beyond the per-waypoint decision. Compared to previous methods, HAM-VLN not only improves various navigation metrics but also reduces the context length by more than 65%. Specifically, HAM-VLN achieves 61.0% Success Rate (SR) on VLN-CE R2R, 52.7% SR on VLN-CE RxR, and 79.7% SR on HM3D-v2 ObjectNav without any training.
Problem

Research questions and friction points this paper is trying to address.

Vision-and-Language Navigation
zero-shot
memory bottleneck
long-horizon navigation
multimodal LLM
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Agentic Memory
Zero-Shot Vision-and-Language Navigation
Persistent World Graph
Context-Efficient Retrieval
Decision-Coupled Memory
πŸ”Ž Similar Papers
A
An Liu
Institute of Automation, Chinese Academy of Sciences, Beijing, China
B
Bingxi Liu
Southern University of Science and Technology, Shenzhen, China; Pengcheng Laboratory, Shenzhen, China
H
Hongyu Ding
Institute of Automation, Chinese Academy of Sciences, Beijing, China; Nanjing University, Nanjing, China
Y
Yixuan Jiang
Institute of Automation, Chinese Academy of Sciences, Beijing, China; Nanjing University, Nanjing, China
Y
Yaran Chen
Xi’an Jiaotong-Liverpool University, Suzhou, China
Fulin Tang
Fulin Tang
Ph.D, University of Chinese Academy of Sciences
SLAM3D reconstructionVLN
C
Cong Leng
MAICRO, Nanjing, China
Hong Zhang
Hong Zhang
School of Cybersecurity and Computer Science, Hebei University
Big DataEdge ComputingInformation SecurityAI
Jian Cheng
Jian Cheng
Beijing, China
computational fluid dynamicshigh-order methodsdiscontinuous Galerkin method