GuidedAttention: Interpretable and Correctable Visual Attention for OOD-Robust Robot Manipulation via Imitation Learning

📅 2026-07-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited interpretability and correctability of end-to-end visuomotor policies in out-of-distribution (OOD) scenarios, which hinder effective human intervention. To overcome this, we propose GuidedAttention, a novel framework that introduces an explicit visual attention mechanism—correctable by the user once prior to execution and automatically propagated throughout the task by a tracking module. By predicting task-relevant keypoints to guide a diffusion-based policy in action generation, our approach maintains end-to-end trainability while significantly enhancing robustness and human–robot collaboration under OOD conditions. Experimental results demonstrate consistent and substantial improvements over baseline methods in both simulation and real-world environments, with particularly strong performance under distribution shifts in object pose and appearance.
📝 Abstract
End-to-end visuomotor policies provide little opportunity for humans to understand or correct the policy's visual attention. We propose GuidedAttention, a visuomotor imitation learning framework that introduces interpretable and correctable visual attention as an explicit intermediate representation. Task-relevant attention keypoints are predicted from camera images and condition a diffusion-based action policy. Users can inspect and optionally correct selected keypoints once at rollout initialization, after which the corrected attention is automatically propagated throughout execution by a tracking module. Experiments in simulation and the real world demonstrate that GuidedAttention consistently improves robot manipulation performance, particularly under positional and appearance out-of-distribution (OOD) conditions.
Problem

Research questions and friction points this paper is trying to address.

visual attention
interpretability
correctability
out-of-distribution
robot manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

interpretable attention
correctable visual attention
imitation learning
diffusion policy
out-of-distribution robustness
Masaki Murooka
Masaki Murooka
National Institute of Advanced Industrial Science and Technology
Robotics
R
Ryoichi Nakajo
Artificial Intelligence Research Center, National Institute of Advanced Industrial Science and Technology (AIST), 2-3-26 Aomi, Koto-ku, Tokyo 135-0064, Japan
Keisuke Shirai
Keisuke Shirai
AIST
Natural Language ProcessingRobotics
Tomohiro Motoda
Tomohiro Motoda
National Institute of Advanced Industrial Science and Technology (AIST)
Robotic manipulationdeep learning
Hanbit Oh
Hanbit Oh
National Institute of Advanced Industrial Science and Technology (AIST)
Robot learningImitation learningLearning from demonstration
R
Ryo Hanai
Artificial Intelligence Research Center, National Institute of Advanced Industrial Science and Technology (AIST), 2-3-26 Aomi, Koto-ku, Tokyo 135-0064, Japan
Yukiyasu Domae
Yukiyasu Domae
AIST
Machine visionManipulationAutomationExperiential autonomy