π€ AI Summary
This study addresses the limited generalization of Vision-Language-Action (VLA) models, which are prone to visual shortcut learning driven by environment-specific features due to insufficient training data diversity. To mitigate this issue, we propose a task-scrubbing domain-adversarial training framework. By revealing varying sensitivities to shortcuts across different backbone architectures, our approach innovatively introduces a task-scrubbing mechanism and a rollback-free action margin metric. Combined with representation-level metric analysis, these components effectively suppress visual shortcuts and enhance the modelβs alignment with and attention to language instructions. Extensive experiments in both simulated and real-world settings demonstrate that the proposed method successfully eliminates visual shortcut phenomena in VLA models, yielding significant improvements in out-of-distribution robustness.
π Abstract
Shortcut learning is a prevalent issue in robot learning. The limited diversity of robot demonstration datasets can mislead policies into exploiting spurious correlations between tasks and irrelevant features, such as viewpoint or background. Collecting sufficiently diverse robot demonstrations is costly and inefficient, motivating algorithmic alternatives. We focus on vision-language-action (VLA) models and discover that different vision-language model backbones exhibit substantially different levels of susceptibility to visual shortcut learning. We find that model behavior correlates with our proposed representation-level metric, action margin, which requires no policy rollouts. Visual shortcuts consistently enter action representations in early layers, with models differing in the extent to which later layers correct them by incorporating language information. To boost models' attention to language, we introduce task scrubbing, a new domain-adversarial training method that decreases models' likelihood of using visual shortcuts and improves VLAs' generalization. Experiments in both simulation and the real world across multiple VLAs and visual cues show that task scrubbing improves out-of-distribution robustness and often eliminates visual shortcut learning.