🤖 AI Summary
This study investigates whether Vision-Language-Action (VLA) models rely on linguistic instructions or solely on visual cues when generating actions. Grounded in a mechanistic interpretability framework, this work employs activation and attribution patching, residual stream analysis, and the LIBERO benchmark to systematically quantify, for the first time, the language sensitivity and grounding mechanisms of π₀.₅ and GR00T N1.7. The findings reveal that both models are highly sensitive to perturbations in directional terms yet remain robust to abstract semantic paraphrasing, while exhibiting significant differences in the distribution of their critical causal layers. By uncovering the internal causal mechanisms underlying distinct VLA architectures, this research provides novel empirical evidence for understanding language dependence in multimodal models.
📝 Abstract
Vision-Language-Action models are designed to generalise across environments and task descriptions, raising the question of whether their action generation actually depends on the language instruction, or whether they largely rely on visual cues and superficial correlations. Robustness to variance in the visual and linguistic observation space is critical for real-world deployment, yet VLAs lack explicit grounding modules and instead rely on the intrinsic language grounding capabilities of their Vision-Language model backbones. For this reason, we conduct a controlled mechanistic interpretability study on the language grounding capabilities of two state-of-the-art Vision-Language-Action models, $π_{0.5}$ and GR00T N1.7, by applying activation and attribution patching to the residual stream of the action generation modules. We systematically corrupt the task instruction of input samples of the LIBERO benchmark following five strategies: synonym replacement, semantic scaling, directional corruption, random object substitution, and empty string. Our experiments find that both models are comparatively insensitive to abstract rephrasing and to referencing non-existent objects, but react strongly to empty task descriptions and, especially, to directional language. During action generation, this sensitivity is concentrated in different loci for each model: mainly in the early, periodic cross-attention layers for GR00T N1.7, versus distributed across the earliest and selected later layers for $π_{0.5}$. For GR00T N1.7, directional perturbations drive some of the largest causal effects while leaving the internal representational geometry comparatively unchanged, a dissociation we do not observe clearly for $π_{0.5}$. Finally, the reliability of attribution patching is model-dependent: it closely tracks activation patching for GR00T N1.7 but not for $π_{0.5}$.