When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limited generalization of Vision-Language-Action (VLA) models, which are prone to visual shortcut learning driven by environment-specific features due to insufficient training data diversity. To mitigate this issue, we propose a task-scrubbing domain-adversarial training framework. By revealing varying sensitivities to shortcuts across different backbone architectures, our approach innovatively introduces a task-scrubbing mechanism and a rollback-free action margin metric. Combined with representation-level metric analysis, these components effectively suppress visual shortcuts and enhance the model’s alignment with and attention to language instructions. Extensive experiments in both simulated and real-world settings demonstrate that the proposed method successfully eliminates visual shortcut phenomena in VLA models, yielding significant improvements in out-of-distribution robustness.
πŸ“ Abstract
Shortcut learning is a prevalent issue in robot learning. The limited diversity of robot demonstration datasets can mislead policies into exploiting spurious correlations between tasks and irrelevant features, such as viewpoint or background. Collecting sufficiently diverse robot demonstrations is costly and inefficient, motivating algorithmic alternatives. We focus on vision-language-action (VLA) models and discover that different vision-language model backbones exhibit substantially different levels of susceptibility to visual shortcut learning. We find that model behavior correlates with our proposed representation-level metric, action margin, which requires no policy rollouts. Visual shortcuts consistently enter action representations in early layers, with models differing in the extent to which later layers correct them by incorporating language information. To boost models' attention to language, we introduce task scrubbing, a new domain-adversarial training method that decreases models' likelihood of using visual shortcuts and improves VLAs' generalization. Experiments in both simulation and the real world across multiple VLAs and visual cues show that task scrubbing improves out-of-distribution robustness and often eliminates visual shortcut learning.
Problem

Research questions and friction points this paper is trying to address.

Shortcut learning
Vision-language-action models
Spurious correlations
Out-of-distribution robustness
Robot learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Shortcut Learning
Task Scrubbing
Domain-Adversarial Training
Action Margin
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
J
Jasper Gerigk
University of Toronto, Toronto, ON M5S 3H5, Canada; Vector Institute, Toronto, ON M5G 0C6, Canada
K
Kenzo Aspuru-Takata
University of Toronto, Toronto, ON M5S 3H5, Canada
Chin-Hsuan Wu
Chin-Hsuan Wu
University of Toronto
Computer VisionRoboticsMachine Learning
Mohammad Mohammadi
Mohammad Mohammadi
Computer Science Student at the University of Toronto
Computer VisionRobotics
S
Shuhong Zheng
University of Toronto, Toronto, ON M5S 3H5, Canada; Vector Institute, Toronto, ON M5G 0C6, Canada
Igor Gilitschenski
Igor Gilitschenski
Assistant Professor, University of Toronto
RoboticsMachine LearningComputer Vision