🤖 AI Summary
This study addresses the challenges of expert data scarcity and low sample efficiency in imitation learning for contact-rich manipulation tasks by proposing the VISTA framework. This method introduces workspace-level equivariance into vision-tactile multimodal fusion, constructing an equivariant diffusion policy through spherical token projection, permutation-equivariant networks, and spherical harmonic representations to achieve spatially consistent action prediction. Both simulation and real-world robotic experiments demonstrate that VISTA substantially improves data efficiency and outperforms existing strong baselines.
📝 Abstract
Imitation learning for contact-rich manipulation requires high-quality expert data that is expensive to obtain. This makes learning a sample-efficient policy a key issue. To address this, we propose VISTA, a workspace-level equivariant visuotactile diffusion policy for data-efficient contact-rich imitation learning. VISTA projects visual and tactile observations into spherical tokens, injects tactile contact cues into visual spherical directions through permutation-equivariant spherical fusion, and rotates the fused harmonic representation using the end-effector orientation. The resulting representation conditions an equivariant diffusion policy to predict spatially consistent actions. Extensive experiments in both simulation and real-world robotic settings show that VISTA substantially improves data efficiency over strong visuotactile imitation learning baselines. Project website: https://vista-paper.github.io/