🤖 AI Summary
This study addresses the limitation of existing Vision-Language-Action (VLA) models in performing high-precision robotic manipulation tasks due to distracted visual attention. To overcome this, we propose a self-supervised spatial grounding technique based on counterfactual visual intervention. Without requiring external annotations, our method directly derives spatial supervision signals from action targets, guiding the model to focus on task-critical regions and achieve human-like precise gaze behavior. We evaluate the proposed approach across four high-precision real-world robot experiments. The results demonstrate that our method significantly enhances attention concentration, comprehensively outperforming both baseline VLA models and mainstream visual grounding approaches.
📝 Abstract
Current Vision-Language-Action (VLA) models often struggle with high-precision robotic manipulation. We attribute this limitation primarily to their visual attention being dispersed across task-irrelevant regions. To address this issue, we propose ActGaze, a training approach that guides VLA policies to gaze on task-relevant regions, much like humans gaze on critical visual cues while executing precise movements. Unlike prior methods that rely on external labels for gaze supervision, ActGaze derives spatial supervision directly from the VLA's own action objective by using counterfactual visual interventions to identify regions that are critical for action prediction. Extensive real-robot experiments on four high-precision robotic manipulation tasks demonstrate that ActGaze induces more focused visual attention on task-relevant regions and consistently outperforms the base VLA policy and other visual-grounding approaches.