🤖 AI Summary
This study addresses the lack of geometric and progress guidance in standard supervision for vision-and-language navigation, as well as the prohibitive overhead of introducing geometric modules during inference. To this end, we propose a privileged spatial guidance framework that operates exclusively during training. The core innovation involves employing a frozen geometric foundation model to provide multi-level spatial priors, combined with hierarchical state supervision, relative heading, and expert route progress objectives to shape navigation representations. Crucially, all auxiliary components are entirely removed at deployment, preserving the original efficient inference pathway. Without requiring additional geometric encoders or data, the proposed method achieves 56.3% SR and 51.4% SPL on R2R-CE, and 54.3% SR on RxR-CE, effectively unifying structure-aware representation learning with efficient deployment.
📝 Abstract
Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. Incorporating depth estimators, explicit maps, point clouds, or geometry foundation models at inference can provide such structure but introduces additional computation, memory overhead, and architectural dependence during deployment. We introduce StageVLN, a training framework that shapes navigation representations through privileged spatial and trajectory guidance while preserving the original inference pathway. A frozen geometry foundation model provides multi-level spatial guidance to hierarchical navigator states, while relative-heading and expert-route progress objectives provide complementary trajectory-state supervision. All auxiliary components are used only during training and removed at deployment. On R2R-CE validation-unseen, StageVLN achieves 56.3\% SR and 51.4\% SPL with a 4B-parameter backbone, without an additional geometry encoder at inference. On RxR-CE, it achieves 54.3\% SR without additional navigation training data or a geometry encoder at inference.