π€ AI Summary
This study addresses the limitations of existing dexterous hand teleoperation systems, which rely heavily on vision and are often constrained to parallel grippers, resulting in low-quality demonstration data and kinematic mismatches during human-to-robot handover. We propose a hand-agnostic, universal dexterous teleoperation interface capable of driving diverse commercial robotic hands through joint-level haptic feedback and contact event rendering. The core innovation lies in a novel reverse teleoperation mechanism that leverages robot backdrivability to pre-align the operatorβs fingers to target configurations, thereby eliminating postural discrepancies upon takeover. Experimental results demonstrate that the proposed approach significantly outperforms existing commercial glove-based solutions in demonstration data quality, data collection throughput, and human-robot collaboration effectiveness for contact-rich manipulation tasks.
π Abstract
Teleoperated demonstrations are a primary source of data for robot manipulation, and teleoperated interventions are a primary mechanism for correcting policies at deployment. Yet most teleoperation systems close the loop through vision alone and are built around parallel-jaw grippers, limiting both what the robot can execute and what the operator can express through it. This is most damaging in shared autonomy, where the operator sees the scene only through occluded cameras and must take over a dexterous hand mid-task, often with an object already grasped. We present DITTO-X, a hand-agnostic dexterous teleoperation interface that renders joint-level force and fingertip contact events from sensing already on the robot hand, and drives three commercial dexterous hands (Sharpa, Wuji, and Inspire) without per-hand redesign. Because the exoskeleton is actuated, DITTO-X also supports reverse teleoperation, in which the robot back-drives the operator's fingers into its own configuration before control is transferred, so the human enters the loop already matched to the state they inherit. Our results show that DITTO-X improves demonstration quality and throughput over a commercial hand-tracking glove, both in regular data collection and in human intervention during policy deployment for contact-rich manipulation tasks. More information can be found from our website: https://tml.stanford.edu/ditto-x/.