🤖 AI Summary
This study addresses the bottleneck of multimodal large models in leveraging historical visual experience for spatial reasoning during closed-loop navigation by proposing a novel paradigm termed video-contextualized navigation. We construct an evaluation benchmark based on Matterport3D and design MV-DualVLN, a dual-view joint planning baseline that decouples goal parsing errors from path execution errors. Experimental results demonstrate that current large models exhibit limited performance on this task, revealing a significant gap between goal recognition and navigation execution. This work provides a critical benchmark and methodological foundation for evaluating and enhancing the instruction-following navigation capabilities of embodied agents that integrate historical video context with real-time observations.
📝 Abstract
Spatial reasoning is fundamental to embodied agents, yet it remains unclear whether spatial understanding can be carried forward to guide sequential interactions. Existing spatial-reasoning benchmarks typically terminate at offline predictions, while navigation benchmarks evaluate spatial reasoning as part of instruction following and exploration. We introduce VCN-Bench, a \textbf{V}ideo-\textbf{C}ontextualized \textbf{N}avigation benchmark for probing closed-loop spatial reasoning over prior visual experience in MLLMs. Given a prior video covering both the initial location and destination, the agent is tasked with reasoning out the instruction-specified target and navigating toward it with the inferred spatial context. Built on Matterport3D, VCN-Bench contains five instruction types, 100k training episodes, and 1,250 evaluation episodes. Navigation serves as the primary evaluation, while diagnostic goal identification helps distinguish destination-resolution errors from subsequent navigation failures. We further propose MV-DualVLN, a planning-oriented baseline that jointly leverages prior video and in-episode observations. Experiments reveal limited navigation performance, a substantial destination-resolution-to-navigation gap, and frequent navigation failures even after correct destination identification.