🤖 AI Summary
This study addresses the limitation of existing agent navigation systems in translating long-term goals into coherent local decisions, particularly their lack of foresight regarding action consequences and future states. To this end, we propose PreAct-Nav, a framework that equips agents with anticipatory reasoning and error-correction capabilities under a frozen policy. Specifically, the method constructs a predictive world sandbox using an Action-Conditional World Model (AC-WM), leverages Vision-Language Models (VLMs) to anchor mid-range subgoals for reasoning, and introduces a persistent memory module to enable dynamic updates. Experimental results demonstrate that PreAct-Nav significantly improves action selection accuracy in long-distance, multi-turn scenarios and enhances navigation robustness within complex urban environments.
📝 Abstract
Urban navigation requires embodied agents to pursue long-horizon goals through local decisions based on egocentric observations. However, existing agentic navigation methods often struggle to translate distant goals into coherent local decisions in large-scale physical environments. Their reliance on linguistic reasoning over transient observations or limited history constrains anticipation of the consequences of actions and future conditions, despite the importance of such foresight for navigating long and complex urban routes. To bridge this gap, we propose PreAct-Nav, an agentic navigation framework that equips frozen policies with anticipatory reasoning for robust urban navigation. Our central idea is to anchor local decisions in persistent medium-horizon subgoals, assess the consequences of predicted actions before execution, and continually update the reasoning context using actual outcomes. At its core, a navigation memory module maintains the active subgoal and relevant experience across decisions, translating distant goals into actionable intermediate objectives. We further introduce a predictive world sandbox that uses an action-conditioned world model (AC-WM) to forecast world dynamics conditioned on candidate movements. A vision-language model (VLM) reasoner interprets these predictions under the current subgoal to retain or revise actions. After execution, real observations are used to assess outcomes, correct inconsistent assumptions, and update memory to continue or reformulate the subgoal. Extensive evaluations demonstrate that the proposed PreAct-Nav improves action selection through memory updates and visual prediction, with more pronounced gains on longer routes and routes with more turns.