Encoded but Not in Control: Revealing the Grounding Gap in Vision-Language Robot Policies

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study identifies a “grounding gap” in Vision-Language-Action (VLA) models, wherein instruction encoding is decoupled from action control. We propose a diagnostic framework based on scene-preserving instruction intervention, integrating interpretability techniques—including linear probing, attention analysis, and inter-layer action lenses—to systematically dissect internal model mechanisms. Our analysis reveals that VLA policies are predominantly governed by scene priors rather than driven by linguistic instructions, exposing latent failures concealed beneath nominal successes. Although language information is encoded, it fails to effectively guide action generation. Accordingly, we establish responsiveness to intent changes as a core evaluation criterion, providing a definitive benchmark for the reliable improvement of VLA models.
📝 Abstract
Instruction following is central to language-conditioned robot policies: language should determine what to do when the same scene permits multiple valid actions. Yet successful execution alone cannot establish whether a policy follows the instruction or infers the task from the scene. We study this ambiguity through scene-preserving instruction interventions, using valid target substitutions, arbitrary nouns, and unrelated sentences while holding the scene fixed. We evaluate vision-language-action (VLA) policies and world-action models (WAMs) in simulation and in real-world experiments. Our analysis addresses three questions: (a) Does task success imply instruction following? When instructions request a different visible object, all evaluated policies predominantly approach and pick up the incorrect original target associated with the scene. (b) Is this failure caused by language insensitivity? Instruction perturbations affect task performance. A layerwise action lens shows intermediate action predictions respond to these perturbations. Linear probes accurately recover instructed targets, indicating modified instructions are encoded despite rarely determining target selection. (c) Why does encoded language fail to control action? Attention analysis indicates weak instruction-token contributions to action generation. Target-token attention can remain focused on the original object, revealing a mismatch between target encoding and visual grounding. UMAP and shared non-negative matrix factorization show target information remains accessible within representations increasingly organized by scene identity. Our findings expose a grounding gap concealed by nominal success and provide a diagnostic framework. They further establish a concrete criterion for progress: policies should reliably follow valid changes in user intent, even when they conflict with scene-favored behavior.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action policies
Instruction following
Grounding gap
Visual grounding
Language control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Policies
Grounding Gap
Instruction Intervention
Attention Analysis
Diagnostic Framework