🤖 AI Summary
This study addresses the significant domain gap in multimodal perception during joint training on human and robot demonstration data, which hinders dexterous manipulation policy learning. To bridge this gap, we propose a framework for constructing a unified multimodal input space. Visually, we introduce a novel approach that erases the human hand from stereo views and replaces it with a photorealistic, noise-augmented robot mesh. Tactilely, we achieve direct signal-space alignment between capacitive gloves and robot fingertip sensors. At the policy level, cross-source co-training is implemented via a diffusion Transformer architecture. Evaluations on three high-precision force-control tasks, including LEGO assembly, demonstrate that our method significantly outperforms policies trained exclusively on robot data, validating the critical role of visual-tactile alignment in enhancing cross-source policy performance.
📝 Abstract
Human demonstrations are a cheap source of data for dexterous manipulation, but co-training a robot policy on them requires closing the human--robot gap in every modality the policy consumes. We present VisTacAlign, a framework for co-training 3D-visual-tactile dexterous policies on human and robot demonstrations. Glove-tracked human hand motion is retargeted to a 17-DoF tactile robot hand with a one-time fingertip correction. The human hand is then erased from both stereo views and replaced by a posed robot-hand mesh painted with pixels from robot recordings, and a real-time stereo foundation model is re-run on the composite, so the human point clouds carry the same stereo errors and visibility as the robot ones. Finally, a capacitive tactile glove is aligned to the robot's fingertip sensors in its signal space, giving one interpretable per-finger force representation. A diffusion transformer consumes point-cloud, proprioceptive, and per-finger tactile tokens. On three real-world tasks requiring precise force -- Lego assembly, plucking strawberries of varying size, and activating and lifting a power drill -- adding aligned human demonstrations to existing robot data improves over robot-only policies, and ablations show that both tactile input and visual alignment are necessary. Project page: https://vis-tac-align.github.io