Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scalability bottleneck in collecting tactile data for dexterous robotic manipulation by proposing an extensible supervised learning paradigm based on human demonstrations. Leveraging a custom-developed tactile motion capture system, we construct the UVTA dataset and design a unified architecture for joint visual-tactile-action prediction, enabling cross-embodiment transfer of physical knowledge. This approach overcomes the data limitations inherent in conventional teleoperation. Evaluated across five manipulation tasks, our method achieves an average success rate of 70%, significantly outperforming existing baselines. Furthermore, performance scales consistently with increasing volumes of human demonstration data, offering an effective pathway toward generalizing tactile skills in embodied intelligence.
📝 Abstract
Dexterous manipulation requires tactile feedback.However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated by a simple premise: hands can change, but the underlying physics of interaction does not. We leverage human tactile data to improve dexterous manipulation policies. Specifically, we first build a tactile motion-capture system that synchronously records images, tactile signals, and hand motions. Using this system, we construct the UVTA dataset spanning five contact-rich tasks, with 1,000 human demonstrations covering diverse interaction patterns and 150 robot demonstrations per task. To transfer the underlying physics of human interaction to robot control, we propose a Unified Visual-Tactile-Action Model that maps both embodiments into aligned tactile and action representations and jointly predicts future action and tactile trajectories. The joint objective enables human demonstrations to supervise contact-aware representation learning, while only robot actions are executed during deployment. In real-robot evaluations across five tasks, our method achieves an average success rate of 70%, outperforming the strongest visual-tactile baseline, which achieves 29%, and an architecture ablation, which achieves 42%. Performance improves consistently with additional human demonstrations and exhibits no saturation at 1,000 demonstrations per task, validating the effectiveness of scalable human tactile data for dexterous manipulation. Project page is available at https://uni-vta.github.io/.
Problem

Research questions and friction points this paper is trying to address.

dexterous manipulation
tactile feedback
human demonstrations
sim-to-real transfer
scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dexterous Manipulation
Visual-Tactile-Action Model
Human Demonstrations
Tactile Motion Capture
Cross-embodiment Transfer
🔎 Similar Papers
No similar papers found.