What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过干预度量方法探讨基于VLM的视觉-语言导航模型如何利用视觉观察、指令和视觉记忆进行决策,并展示这些模型能够编码导航进度并保留语义结构。
📝 Abstract
Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While this fusion of language instructions and visual observations allows multimodal reasoning, it obscures how information is routed across modalities or what mechanisms drive navigation decisions. Thus, it remains unclear whether VLN models ground their predictions in relevant semantic cues or can track task progress. In this work, we study the interpretability and steerability of VLN models. We use intervention-based metrics that measure how visual observations, instructions, and visual memory causally influence navigation decisions. Our results show that these navigation policies are sensitive to all input modalities and do not depend on a single one. We further show that these agents encode navigation progress and retain semantic structure from their VLM backbones, enabling concept-level steering through internal activations. Finally, we extract activation vectors for abstract behaviors to transfer them zero-shot to out-of-distribution real-world scenarios, improving performance without additional fine-tuning.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Navigation
Multimodal Reasoning
Interpretability
Steerability
Causal Influence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Navigation
Interpretability
Steerability
Zero-shot Transfer
Activation Vectors