🤖 AI Summary
This study addresses the challenges in dexterous manipulation retargeting arising from kinematic discrepancies between human and robotic hands, which often lead to contact pattern loss and inefficient single-trajectory training. To overcome these limitations, this work proposes a novel tactile-guided reinforcement learning framework that leverages glove-based tactile signals to direct policy search, thereby preserving human contact patterns. Furthermore, it jointly trains a single universal retargeter across multiple trajectories and employs knowledge distillation to derive a deployment controller independent of tactile sensing. Experimental evaluations on pen and hammer manipulation tasks demonstrate successful transfer of natural contact behaviors. The proposed approach exhibits significantly superior generalization compared to baseline methods and single-trajectory optimization schemes, achieving efficient and robust transfer of dexterous manipulation behaviors.
📝 Abstract
Glove-based motion capture is emerging as a scalable approach to collecting dexterous-hand demonstration data. However, due to the kinematic gap between the human and robot hand, the recorded human motions cannot be executed directly on the robot, especially for contact-rich tool-use tasks involving in-hand reorientation. Prior work bridges this gap in simulation through reinforcement learning (RL) or trajectory optimization, but the human contact pattern is hard to preserve under such formulations, often producing unnatural manipulation and unstable functional grasps. These methods also train a separate policy or solve a separate optimization for each reference trajectory, which is inefficient. To solve these problems, we propose DexTaG, a tactile-guided RL framework for dexterous manipulation. During training, tactile signals captured by the glove guide policy search toward the measured human contact pattern, reducing reliance on precise reference geometry for contact supervision. To improve efficiency, we train a single generalizable retargeter jointly on all training trajectories of the same object. The retargeter is further distilled into a tactile-free student controller conditioned on the target object trajectory for real-world deployment. On marker-pen and hammer manipulation tasks, DexTaG learns natural, contact-rich behaviors that baselines with distance-based contact heuristics fail to learn, generalizes to held-out trajectories of the same object and task, and outperforms single-trajectory baselines on OakInk2.