Geometry-Preserving Human-to-Robot Upper-Body Motion Retargeting from Monocular Video

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of motion retargeting from monocular videos to robots, where significant kinematic discrepancies between humans and robots complicate the preservation of fine hand details. We propose a geometry-preserving framework that integrates unified body-hand reconstruction with morphology-agnostic geometric transfer. By leveraging differentiable MHR state joint constraints and parameter-space repair mechanisms, combined with multi-stage inverse kinematics and VAM/WAM system modules, our method achieves coordinated upper-body motion mapping from video to dual-arm dexterous robots. Experimental results demonstrate that the hand reprojection error is reduced to 7.21 pixels. All generated trajectories are validated in simulation, and sign language as well as grasping motions are successfully deployed on physical robots.
📝 Abstract
Monocular RGB video provides an accessible source of human demonstrations for upper-body robot motion, yet video-driven human-to-robot transfer remains challenging because body and hand motion are recovered at different spatial scales, human and robot kinematics differ substantially, and fine distal motion is difficult to preserve across embodiments. We present a geometry-preserving motion-retargeting framework that integrates unified body--hand reconstruction with morphology-independent geometric transfer. Frame-wise body estimates, video-level observations, and detailed hand evidence jointly constrain a single differentiable Momentum Human Rig (MHR) state, while transient hand artifacts are repaired in parameter space. The reconstructed motion is represented by arm-segment directions, elbow configuration, relative palm orientation, and bilateral wrist relations, and is realized on the target robot through multi-stage inverse kinematics and robot-specific hand adaptation. Within the broader system, Across-VAM provides video generation, whereas Across-WAM performs human-to-robot motion mapping. The method is evaluated on 16 monocular videos comprising 1,769 source frames, including 10 signing and six reach-to-grasp sequences. Unified reconstruction reduces mean hand reprojection error from 22.36 to 7.21 pixels relative to SAM 3D Body. All 16 retargeted trajectories completed kinematic simulation playback, and representative signing and reach-to-grasp motions were further demonstrated on a physical robot. The results demonstrate a unified pipeline from monocular human video to coordinated upper-body motion on a dual-arm dexterous robot.
Problem

Research questions and friction points this paper is trying to address.

Motion Retargeting
Monocular Video
Human-to-Robot Transfer
Upper-Body Motion
Geometry-Preserving
Innovation

Methods, ideas, or system contributions that make the work stand out.

Motion Retargeting
Monocular Video
Differentiable Human Rig
Inverse Kinematics
Unified Body-Hand Reconstruction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Xiaoyu Yang
Xiaoyu Yang
University of Cambridge
Speech recognitionmachine learning
S
Sen Han
Beijing SFlare Robotics Technology Co., Ltd., Beijing, China
Da Li
Da Li
Beijing institute of technology
Radar systemCross-modal learningSensor fusion
N
Nan Wu
Across Physics, Beijing, China