motion retargeting

Algorithms and modeling approaches that transfer or expand motion demonstrations to different characters or embodiments while preserving timing, contacts, and coordination. This includes generating whole-body trajectories from single demonstrations, rendering across humanoid models and viewpoints, and extracting geometric supervision (e.g., future end-effector waypoints) from source videos.

motionretargeting

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of existing motion transfer methods, which rely on predefined human skeletal structures and annotated data, thereby struggling to generalize across species. To overcome these constraints, the authors propose Motion4Motion—a training-free, cross-species motion transfer framework that eschews explicit skeleton modeling in favor of optical flow representations derived directly from video. By aligning motion features across domains during inference, the method enables high-quality motion transfer between diverse subjects—including across species—without requiring retraining or fine-tuning. This approach significantly enhances generalization capability and practical flexibility, outperforming current baselines across a range of characters and unlocking novel applications in animation production and beyond.

cross-speciesdiverse charactersmotion transfer

Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D Motions

Dec 09, 2025
JL
Jianan Li
🏛️ The Chinese University of Hong Kong | Shanghai AI Laboratory | Shanghai Jiao Tong University | Monash University

This paper addresses the challenge of learning physically plausible 3D character controllers directly from monocular 2D video keypoints—without requiring scarce 3D motion-capture data or pre-trained motion reconstruction models. Methodologically, it introduces Mimic2DM, a hierarchical control framework: an upper-level Transformer autoregressively generates diverse 2D action sequences, while a lower-level module implicitly models 3D motion via multi-view aggregation and jointly optimizes reprojection loss and physics-based simulation rewards for end-to-end training. Its key contribution is the first demonstration that a physically consistent and generalizable 3D control policy can be trained using *only* 2D keypoint supervision. Experiments on complex human-object interaction (HOI) and non-human character tasks—including dance, soccer dribbling, and animal locomotion—validate superior motion diversity, visual realism, and physical plausibility. Mimic2DM significantly improves the practicality and robustness of existing 2D-driven 3D control approaches.

Generating diverse, physically plausible motions without 3D dataLearning 3D character control from 2D video dataOvercoming limitations of 3D motion reconstruction methods

This work addresses the limited generalizability of existing cross-embodiment video generation methods, which suffer from entangled motion and morphology representations and rely on paired data for target embodiments. To overcome these limitations, the authors propose a motion-morphology disentangled modeling framework that enables rapid adaptation to new robots without requiring paired data, leveraging a shared motion model and lightweight embodiment adapters. A novel branch-isolated attention mechanism is introduced to effectively separate motion conditioning from embodiment-specific modulation. The study also presents the first large-scale synthetic dataset of cross-embodiment paired videos. Experimental results demonstrate high motion fidelity and embodiment consistency on both synthetic and real-world benchmarks, with successful zero-shot transfer to unseen humanoid embodiments without retraining the shared motion model.

cross-embodimentembodiment adaptationmotion transfer

This work addresses the challenge of generalizing imitation-learning-based manipulation policies across objects with significant geometric discrepancies—e.g., pouring liquid into unseen containers with novel shapes and poses. We propose the Motion Transfer Frame (MTF) framework, which automatically identifies geometry-agnostic key points and dynamic reference frames grounded in both object geometry and task semantics. MTF integrates geometric-aware keypoint localization, reference-frame binding, and kinematic constraint modeling, and supports closed-loop validation from simulation to real robots. Its core contribution is geometry-invariant trajectory transfer that simultaneously enforces critical pose constraints (e.g., cup upright orientation), collision-free motion, and task success. Experiments demonstrate >92% pouring success across diverse unseen container configurations, substantially improving cross-morphology generalization of manipulation skills.

Ensuring motion constraints, collision-free paths, and task success.Generating motion plans using motion transfer frames for arbitrary tasks.Transferring kinesthetic demonstrations across diverse object geometries.

MagicPose4D: Crafting Articulated Models with Appearance and Motion Control

May 22, 2024
HZ
Hao Zhang
🏛️ University of Illinois Urbana-Champaign | University of Southern California

Existing 4D content generation methods rely on text prompts, limiting precise control over complex or rare motions. MagicPose4D addresses this by introducing a novel two-stage 4D generation framework that accepts monocular videos or mesh sequences as motion priors, enabling fine-grained co-modeling of appearance and motion. Its key contributions are: (1) a cross-category motion transfer module; (2) a global-local Chamfer loss combined with kinematic-chain-based skeletal constraints to ensure both geometric fidelity and physical plausibility; and (3) a multi-source supervision strategy integrating 2D image reconstruction, pseudo-3D supervision, dynamic rigid interpolation, and skeleton-driven motion transfer. Experiments demonstrate significant improvements over state-of-the-art methods in motion accuracy, temporal coherence, and cross-category generalization. Notably, MagicPose4D achieves robust motion transfer across categories without fine-tuning.

Enables precise motion control in 4D generation using video or mesh inputsFacilitates cross-category motion transfer with kinematic-chain-based skeletonImproves 4D reconstruction via dual-phase shape and motion extraction

Latest Papers

What's happening recently
View more

Existing approaches struggle to efficiently transfer human motion styles to diverse locomotion content on humanoid robots while ensuring physical feasibility, often producing motions that are kinematically implausible or dynamically unstable. This work proposes a bio-inspired generative-to-control framework that, for the first time, enables reusable style transfer from short human demonstrations without retraining, allowing continuous adjustment of style intensity. The method integrates a physics-aware multi-conditional latent diffusion model with classifier-free guidance, contact consistency constraints, and temporal smoothing regularization. It further introduces a preview-based whole-body tracking scheme and a clustering-distillation training strategy. Evaluated on the Unitree G1 robot, the approach achieves a 96.0% execution success rate, significantly reducing contact penetration and jitter artifacts while enabling high-quality stylized execution across a wide range of motion content.

expressive motionhumanoid robotsmotion generation

Humanoid robots face significant challenges in loco-manipulation tasks due to the scarcity of high-quality demonstration data in high-dimensional action spaces. This work proposes an automatic data synthesis method that leverages a small set of source demonstrations and contact-aware whole-body motion planning to transfer contact-rich skills to novel states. By jointly optimizing locomotion and manipulation, the approach generates diverse, stable, and collision-free whole-body behaviors at scale—marking the first large-scale automated generation of loco-manipulation data for humanoid robots. The synthesized dataset enables cross-object-pose generalization and supports a new simulation benchmark comprising nine distinct tasks. Policies trained with this data achieve a 20% performance gain over those trained solely on real demonstrations, establishing a systematic foundation for studying data generation and visuomotor policy learning.

data generationhumanoid robotsimitation learning

This work addresses the limitations in humanoid robot skill learning imposed by the high cost of acquiring real human demonstration data, insufficient motion diversity, and inadequate coverage of individual variability. The authors propose a purely synthetic learning framework that eliminates the need for real human demonstrations by leveraging generative AI to translate textual prompts into diverse and realistic human motion sequences. These synthesized motions are embedded into virtual video scenes to train robots within a simulation environment through a combination of reinforcement and imitation learning. This approach achieves, for the first time, fully synthetic-data-driven acquisition of diverse motor skills. Evaluated across four simulated tasks, the method not only successfully accomplishes target objectives but also demonstrates strong generalization capabilities to complex motion variations, thereby overcoming reliance on real-world demonstration data.

humanoid robotslearning from demonstrationsmotion diversity

This study investigates how to effectively organize heterogeneous robotic demonstration data to enhance cross-embodiment transfer performance. Through controlled simulation experiments, it systematically compares the efficacy of unpaired large-scale data against structured paired data—such as demonstrations aligned by scene, task, or trajectory—under varying morphologies and viewpoints. The findings reveal that, for morphology differences, structured data analogies are more effective than merely increasing data diversity, highlighting distinct data structure requirements for morphology transfer versus viewpoint transfer. By optimizing data composition alone, the approach achieves an average 22.5% improvement in success rate on real-world cross-embodiment transfer tasks, underscoring the critical role of data analogy in enabling effective embodiment-agnostic skill transfer.

cross-embodiment transferdata analogydata diversity

This work addresses the challenge of acquiring whole-body manipulation demonstrations for humanoid robots, which is hindered by the high cost, skill requirements, and inefficiency of existing teleoperation systems. The authors propose a portable, robot-free data collection framework that leverages lightweight VR hardware and a UMI-inspired gripper to simultaneously capture sparse human body keypoints, wrist-view images, and gripper actions. A high-level policy predicts future keypoints and retargets them into full-body reference commands for a whole-body controller to execute. This approach enables, for the first time, the collection of whole-body manipulation demonstrations without requiring a physical robot, extending the UMI paradigm to full-body behaviors and supporting cross-task skill transfer. The effectiveness and generalizability of the learned skills are validated across five real-world scenarios.

demonstration datahumanoid robotsskill learning

Hot Scholars

MH

Marco Hutter

Professor of Robotics, ETH Zurich
Legged RoboticsRoboticsControl
PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
MH

Ming-Hsuan Yang

University of California at Merced; Google DeepMind
Computer VisionMachine LearningArtificial Intelligence
GS

Guanya Shi

Assistant Professor, CMU RI | Amazon Scholar, FAR (Frontier AI & Robotics)
RoboticsRobot LearningReinforcement LearningControl