π€ AI Summary
This study addresses the vulnerability of imitation learning policies that over-rely on visual modalities and fail under intermittent visual occlusion. To mitigate this, we propose a directed modality dropout method that leverages attention mechanisms to precisely identify and selectively discard dominant modalities, coupled with entropy regularization to enhance policy robustness. This approach overcomes the limitations of conventional random dropout strategies. Experimental evaluations on real-world bimanual robotic tasks demonstrate that our method significantly improves the modelβs adaptability to specific sensory deprivation. Notably, under conditions of visual loss, the proposed approach achieves substantially higher task success rates compared to both standard baselines and random dropout methods, highlighting its effectiveness in ensuring reliable policy execution despite partial sensory degradation.
π Abstract
Imitation learning policies that integrate multiple sensory modalities are prone to overreliance on a dominant modality, such as vision, during training, which can disrupt policy execution when that modality is lost at inference time. In this paper, we introduce Targeted Modality Dropout (TMD), in which the dependence on each modality is estimated using attention and the most dominant modality is selectively dropped. This is combined with entropy regularization over the dependence distribution. Through real-robot evaluation using a bimanual manipulator, we show that under vision loss the success rate of the baseline policy drops substantially, whereas TMD sustains task execution. In contrast, a conventional dropout that selects the dropped modality at random, without the entropy regularization, fails on many tasks even without vision loss.