🤖 AI Summary
This work addresses the limited interpretability and correctability of end-to-end visuomotor policies in out-of-distribution (OOD) scenarios, which hinder effective human intervention. To overcome this, we propose GuidedAttention, a novel framework that introduces an explicit visual attention mechanism—correctable by the user once prior to execution and automatically propagated throughout the task by a tracking module. By predicting task-relevant keypoints to guide a diffusion-based policy in action generation, our approach maintains end-to-end trainability while significantly enhancing robustness and human–robot collaboration under OOD conditions. Experimental results demonstrate consistent and substantial improvements over baseline methods in both simulation and real-world environments, with particularly strong performance under distribution shifts in object pose and appearance.
📝 Abstract
End-to-end visuomotor policies provide little opportunity for humans to understand or correct the policy's visual attention. We propose GuidedAttention, a visuomotor imitation learning framework that introduces interpretable and correctable visual attention as an explicit intermediate representation. Task-relevant attention keypoints are predicted from camera images and condition a diffusion-based action policy. Users can inspect and optionally correct selected keypoints once at rollout initialization, after which the corrected attention is automatically propagated throughout execution by a tracking module. Experiments in simulation and the real world demonstrate that GuidedAttention consistently improves robot manipulation performance, particularly under positional and appearance out-of-distribution (OOD) conditions.