🤖 AI Summary
This study addresses the challenge of mapping human demonstrations to dexterous robotic manipulation, where inverse kinematics neglects dynamics and reinforcement learning or model predictive control (MPC) suffers from low sample efficiency and training instability. To overcome these limitations, this work proposes a flow matching-based generative neural retargeting framework. By assuming feasible trajectories reside on a low-dimensional manifold, the method reformulates retargeting as a conditional sampling problem, integrating a real-time simulation engine with contact force annotation techniques for efficient dynamic mapping. This approach avoids per-trajectory optimization and reduces hyperparameter sensitivity. It achieves a 56.2% success rate using only 8.5% of the samples required by MPC, which attains merely 27.2%. Furthermore, this research constructs a large-scale, high-fidelity dataset comprising 223,000 demonstrations across 3,300 object geometries.
📝 Abstract
Human demonstrations are a scalable data source for learning dexterous manipulation, but the embodiment gap prevents human motion from being executed directly on robots. Inverse kinematics (IK) retargets human motion to robots efficiently but ignores dynamics, often producing infeasible motions. Reinforcement learning (RL) and sampling-based model predictive control (MPC) are commonly employed to yield dynamically feasible motions, but both are sample-inefficient and sensitive to hyperparameters. RL suffers from costly and unstable training and tedious reward engineering; MPC avoids policy optimization, yet retargets each trajectory in isolation, and solving one does not make the next easier. Sampling cost grows rapidly with dataset size and task difficulty. We hypothesize that dynamically feasible trajectories concentrate near a low-dimensional manifold shared across demonstrations, so that retargeting can be reduced to sampling from that manifold, conditioned on human motion, rather than solving a fresh optimization problem for every demonstration. We propose \textbf{Generative Neural Retargeting} (GNR), which uses a flow matching model to sample feasible trajectories. GNR outperforms MPC with only $8.5\%$ of the samples required by MPC, achieving a success rate of $56.20\%$ compared to $27.20\%$ for MPC. GNR can be used for scalable and efficient retargeting of large-scale, long-horizon, and millimeter precision human demonstrations: by applying GNR within a real-to-sim data engine, we produce a dexterous manipulation dataset with dense contact-force labels, spanning $223$k demonstrations and $3.3$k object geometries.