Targeted Modality Dropout for Real-Robot Manipulation Robust to Intermittent Vision Loss

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the vulnerability of imitation learning policies that over-rely on visual modalities and fail under intermittent visual occlusion. To mitigate this, we propose a directed modality dropout method that leverages attention mechanisms to precisely identify and selectively discard dominant modalities, coupled with entropy regularization to enhance policy robustness. This approach overcomes the limitations of conventional random dropout strategies. Experimental evaluations on real-world bimanual robotic tasks demonstrate that our method significantly improves the model’s adaptability to specific sensory deprivation. Notably, under conditions of visual loss, the proposed approach achieves substantially higher task success rates compared to both standard baselines and random dropout methods, highlighting its effectiveness in ensuring reliable policy execution despite partial sensory degradation.
πŸ“ Abstract
Imitation learning policies that integrate multiple sensory modalities are prone to overreliance on a dominant modality, such as vision, during training, which can disrupt policy execution when that modality is lost at inference time. In this paper, we introduce Targeted Modality Dropout (TMD), in which the dependence on each modality is estimated using attention and the most dominant modality is selectively dropped. This is combined with entropy regularization over the dependence distribution. Through real-robot evaluation using a bimanual manipulator, we show that under vision loss the success rate of the baseline policy drops substantially, whereas TMD sustains task execution. In contrast, a conventional dropout that selects the dropped modality at random, without the entropy regularization, fails on many tasks even without vision loss.
Problem

Research questions and friction points this paper is trying to address.

imitation learning
multimodal sensory
modality overreliance
intermittent vision loss
robot manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Targeted Modality Dropout
Imitation Learning
Multimodal Robustness
Attention Mechanism
Entropy Regularization
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
G
Genki Shikada
Fujitsu Limited, Kanagawa 211-8588, Japan; Faculty of Science and Engineering, Waseda University, Tokyo 169-8050, Japan
K
Kazuki Osamura
Fujitsu Limited, Kanagawa 211-8588, Japan
M
Masaru Ide
Fujitsu Limited, Kanagawa 211-8588, Japan
Tetsuya Ogata
Tetsuya Ogata
Professor, Waseda University / Joint-appointed Fellow, AIST / Visiting Professor, NII
Deep Predictive LearningPhysical AIDevelopmental Robotics
Kanata Suzuki
Kanata Suzuki
Fujitsu Limited / Waseda University
Robot Learning