🤖 AI Summary
This study addresses the infeasible motions and error accumulation inherent in traditional motion retargeting caused by morphological discrepancies. To overcome these limitations, this work proposes an end-to-end video-to-robot-action mapping framework that eliminates explicit human representations by directly generating robot commands from monocular videos. Specifically, the method learns a robot-oriented implicit representation to capture cross-morphology structural commonalities and introduces a contact-aware optimization mechanism to ensure physical consistency. Experimental results demonstrate that the proposed framework significantly enhances motion accuracy and robustness, achieving higher execution success rates on real robots and lower inference latency. Ultimately, this approach establishes a new paradigm for cross-morphology motion transfer.
📝 Abstract
Learning executable motions from human videos offers a scalable solution for humanoid robots to acquire demonstration motions. However, existing pipelines typically first construct an explicit human motion representation and then convert it into robot motions via motion retargeting. Although such methods can effectively leverage large volumes of existing human data for training, the substantial differences between humans and humanoid robots in locomotion mechanisms and joint degree-of-freedom configurations make motions generated by this human-representation-centric approach difficult to execute on robots. Furthermore, errors introduced during human motion estimation inevitably propagate to the retargeting stage and cannot be eliminated via joint optimization. We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions. Discarding the explicit human representation, this framework learns robot-oriented implicit representations directly from visual observations, enabling the model to capture cross-morphology motion structures. To generate motions more suitable for robot execution, we further design a contact-aware motion optimization mechanism to improve temporal consistency and physical plausibility. Experiments show that BeyondRetarget significantly improves the accuracy and robustness of generated robot motions, while achieving higher execution success rates and lower latency in both simulation environments and real humanoid robots.