🤖 AI Summary
This work addresses the challenges of high latency, physical infeasibility, and poor robustness to noise and partial observations in online motion retargeting for humanoid robot teleoperation. We propose a retargeting-free, end-to-end deep reinforcement learning framework that leverages vector quantization to construct a whole-body atomic motion primitive codebook. This approach directly maps raw human motions to joint commands while projecting out-of-distribution or partially missing observations onto plausible prototypes to recover complete motions. The framework supports multimodal inputs, including VR, motion capture, text, and monocular video. Simulation and hardware experiments on the Unitree G1 demonstrate that our method significantly outperforms existing baselines in both latency and robustness, achieving stable and efficient whole-body teleoperation control.
📝 Abstract
Humanoid whole-body teleoperation translates human motion into stable robot behavior in real time. Existing systems typically rely on online motion retargeting to bridge human--robot morphological differences, but this process adds latency and can produce physically infeasible targets. Meanwhile, diverse, noisy, and partial human-motion observations often fall outside the training distribution, potentially causing unstable robot behavior. We propose a retargeting-free policy that maps raw human motion directly to robot joint commands in a single forward pass, eliminating online kinematic adaptation. To improve robustness, we learn a codebook of full-body motion primitives that projects out-of-distribution observations onto plausible motion prototypes and recovers full-body motion from partial inputs. Experiments on a Unitree~G1 in simulation and on hardware, using virtual reality, optical mocap, text-to-motion generation, and monocular video inputs, show that our method outperforms baselines in latency and robustness.