Equivariant Visual-Tactile Diffusion Policy for Contact-Rich Manipulation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of expert data scarcity and low sample efficiency in imitation learning for contact-rich manipulation tasks by proposing the VISTA framework. This method introduces workspace-level equivariance into vision-tactile multimodal fusion, constructing an equivariant diffusion policy through spherical token projection, permutation-equivariant networks, and spherical harmonic representations to achieve spatially consistent action prediction. Both simulation and real-world robotic experiments demonstrate that VISTA substantially improves data efficiency and outperforms existing strong baselines.
📝 Abstract
Imitation learning for contact-rich manipulation requires high-quality expert data that is expensive to obtain. This makes learning a sample-efficient policy a key issue. To address this, we propose VISTA, a workspace-level equivariant visuotactile diffusion policy for data-efficient contact-rich imitation learning. VISTA projects visual and tactile observations into spherical tokens, injects tactile contact cues into visual spherical directions through permutation-equivariant spherical fusion, and rotates the fused harmonic representation using the end-effector orientation. The resulting representation conditions an equivariant diffusion policy to predict spatially consistent actions. Extensive experiments in both simulation and real-world robotic settings show that VISTA substantially improves data efficiency over strong visuotactile imitation learning baselines. Project website: https://vista-paper.github.io/
Problem

Research questions and friction points this paper is trying to address.

contact-rich manipulation
imitation learning
sample efficiency
visuotactile
Innovation

Methods, ideas, or system contributions that make the work stand out.

Equivariant Diffusion Policy
Visuotactile Fusion
Spherical Harmonics
Contact-Rich Manipulation
Imitation Learning
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
L
Lik Hang Kenny Wong
Department of Computer Science and Engineering, The Chinese University of Hong Kong
Y
Yiyao Ma
Department of Computer Science and Engineering, The Chinese University of Hong Kong
Xiu-Shen Wei
Xiu-Shen Wei
Professor, Southeast University
Computer VisionMachine LearningArtificial Intelligence
Z
Zelong Tan
Department of Computer Science and Engineering, The Chinese University of Hong Kong
Z
Zhuheng Song
Department of Computer Science and Engineering, The Chinese University of Hong Kong
D
Dongsheng Xie
Department of Computer Science and Engineering, The Chinese University of Hong Kong
K
Kai Chen
Department of Computer Science and Engineering, The Chinese University of Hong Kong
Q
Qi Dou
Department of Computer Science and Engineering, The Chinese University of Hong Kong