VisTacAlign: Co-Training Dexterous Policies on Tactile Human and Robot Demonstrations

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant domain gap in multimodal perception during joint training on human and robot demonstration data, which hinders dexterous manipulation policy learning. To bridge this gap, we propose a framework for constructing a unified multimodal input space. Visually, we introduce a novel approach that erases the human hand from stereo views and replaces it with a photorealistic, noise-augmented robot mesh. Tactilely, we achieve direct signal-space alignment between capacitive gloves and robot fingertip sensors. At the policy level, cross-source co-training is implemented via a diffusion Transformer architecture. Evaluations on three high-precision force-control tasks, including LEGO assembly, demonstrate that our method significantly outperforms policies trained exclusively on robot data, validating the critical role of visual-tactile alignment in enhancing cross-source policy performance.
📝 Abstract
Human demonstrations are a cheap source of data for dexterous manipulation, but co-training a robot policy on them requires closing the human--robot gap in every modality the policy consumes. We present VisTacAlign, a framework for co-training 3D-visual-tactile dexterous policies on human and robot demonstrations. Glove-tracked human hand motion is retargeted to a 17-DoF tactile robot hand with a one-time fingertip correction. The human hand is then erased from both stereo views and replaced by a posed robot-hand mesh painted with pixels from robot recordings, and a real-time stereo foundation model is re-run on the composite, so the human point clouds carry the same stereo errors and visibility as the robot ones. Finally, a capacitive tactile glove is aligned to the robot's fingertip sensors in its signal space, giving one interpretable per-finger force representation. A diffusion transformer consumes point-cloud, proprioceptive, and per-finger tactile tokens. On three real-world tasks requiring precise force -- Lego assembly, plucking strawberries of varying size, and activating and lifting a power drill -- adding aligned human demonstrations to existing robot data improves over robot-only policies, and ablations show that both tactile input and visual alignment are necessary. Project page: https://vis-tac-align.github.io
Problem

Research questions and friction points this paper is trying to address.

dexterous manipulation
human-robot gap
co-training
tactile sensing
visual alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

VisTacAlign
dexterous manipulation
visual-tactile co-training
diffusion transformer
human-robot alignment
💼 Related Jobs
No related jobs found.
J
Julien Poffet
Department of Mechanical and Process Engineering (D-MAVT), ETH Zurich, Zurich, Switzerland. Work done as a visiting student at Stanford University.
M
Matthew Strong
Department of Computer Science, Stanford University, Stanford CA, USA
A
Ankush Dhawan
Department of Mechanical Engineering, Stanford University, Stanford CA, USA
B
Baiyu Shi
Department of Mechanical Engineering, Stanford University, Stanford CA, USA
S
Shalika Neelaveni
Department of Mechanical Engineering, Stanford University, Stanford CA, USA
Y
Yujia Yuan
Department of Electrical Engineering, Stanford University, Stanford CA, USA
Zhenan Bao
Zhenan Bao
Department of Chemical Engineering, Stanford University, Stanford CA, USA
Monroe Kennedy III
Monroe Kennedy III
Assistant Professor of Mechanical Engineering, Stanford University
RoboticsRobotic AssistantsAssistive RoboticsRobotic ManipulationDynamics and Controls