🤖 AI Summary
This study addresses the challenge of decoupled visual and exploratory memory in zero-shot navigation, which impedes effective decision-making. To overcome this, we propose a "memory visualization" paradigm that projects topological graph nodes and exploration states into the first-person perspective, rendering memory directly observable for route selection. This enables Vision-Language Models (VLMs) to jointly evaluate goal relevance and exploration progress without requiring post-processing or additional fusion steps. Our work reveals the critical influence of memory representation on VLM-based decision-making. The proposed approach achieves state-of-the-art performance on HM3D with an 81.2% success rate while requiring only 7.5% of the VLM invocations used by WMNav. Furthermore, real-world robotic experiments validate the practical feasibility of deploying this method.
📝 Abstract
When searching for an object, people choose their next move by considering both likely target locations and places already explored. The current view can cue place-associated memories, bringing target relevance and prior exploration into the same spatial context. In many zero-shot object navigation (ZSON) methods, however, vision-language models (VLMs) infer promising search areas from egocentric images, while exploration history is represented separately, e.g., as text or maps. This separation either requires an additional fusion step or leaves the correspondence between memory and route choices implicit for the VLM to recover. We instead make exploration memory directly visible on visual route choices. We propose MarvisNav, a ZSON framework that maintains a topological graph and projects candidate nodes together with their exploration states onto egocentric views as memory-bearing visual route choices. These states capture local exploration progress beyond binary visitation. By binding exploration state directly to each visual candidate, MarvisNav enables the VLM to jointly evaluate target relevance and exploration state without a separate post-hoc fusion or reranking stage. Without policy training, MarvisNav achieves state-of-the-art performance on HM3D (81.2% SR and 42.5% SPL), while remaining competitive on MP3D. It also outperforms representative VLM-based methods with far fewer VLM calls (e.g., 7.5% of WMNav). Real-robot experiments across diverse scenes further validate its practical deployability. Beyond MarvisNav, our study shows that memory representation shapes VLM decisions and ZSON performance, highlighting that effective memory use depends not only on its availability, but also on how it is represented. Code and project page will be available at \url{https://wangjincheng1998.github.io/MarvisNav/}.