BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the infeasible motions and error accumulation inherent in traditional motion retargeting caused by morphological discrepancies. To overcome these limitations, this work proposes an end-to-end video-to-robot-action mapping framework that eliminates explicit human representations by directly generating robot commands from monocular videos. Specifically, the method learns a robot-oriented implicit representation to capture cross-morphology structural commonalities and introduces a contact-aware optimization mechanism to ensure physical consistency. Experimental results demonstrate that the proposed framework significantly enhances motion accuracy and robustness, achieving higher execution success rates on real robots and lower inference latency. Ultimately, this approach establishes a new paradigm for cross-morphology motion transfer.
📝 Abstract
Learning executable motions from human videos offers a scalable solution for humanoid robots to acquire demonstration motions. However, existing pipelines typically first construct an explicit human motion representation and then convert it into robot motions via motion retargeting. Although such methods can effectively leverage large volumes of existing human data for training, the substantial differences between humans and humanoid robots in locomotion mechanisms and joint degree-of-freedom configurations make motions generated by this human-representation-centric approach difficult to execute on robots. Furthermore, errors introduced during human motion estimation inevitably propagate to the retargeting stage and cannot be eliminated via joint optimization. We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions. Discarding the explicit human representation, this framework learns robot-oriented implicit representations directly from visual observations, enabling the model to capture cross-morphology motion structures. To generate motions more suitable for robot execution, we further design a contact-aware motion optimization mechanism to improve temporal consistency and physical plausibility. Experiments show that BeyondRetarget significantly improves the accuracy and robustness of generated robot motions, while achieving higher execution success rates and lower latency in both simulation environments and real humanoid robots.
Problem

Research questions and friction points this paper is trying to address.

humanoid robots
motion retargeting
monocular video
executable motions
cross-morphology
Innovation

Methods, ideas, or system contributions that make the work stand out.

End-to-end framework
Implicit representation
Monocular video
Contact-aware optimization
Motion retargeting
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tianyu Xiong
School of Electronic Science and Engineering, Nanjing University, Nanjing, China
Yi Lu
Yi Lu
Professor of Electrical and Computer Engineering, University of Illinois
Cloud computingnetwork algorithmsperformance evaluation
J
Jinrui Wang
School of Electronic Science and Engineering, Nanjing University, Nanjing, China
Ziqi Liang
Ziqi Liang
Ant Group
Text-to-SpeechVoice ConverisonLLMRL
D
Dandan Lei
Jiangsu Mobile Information System Integration Co., Ltd., Nanjing, China
X
Xiaoyang Zhou
China Mobile Zijin (Jiangsu) Innovation Research Institute Co., Ltd., Nanjing, China
X
Xiao-xiao Long
School of Intelligence Science and Technology, Nanjing University, Suzhou, China
Qiu Shen
Qiu Shen
Nanjing University
Xun Cao
Xun Cao
Nanjing University
Computational PhotographyComputational ImagingImage & Video Processing