procedural animation

Algorithmic generation of motion sequences from models, constraints, and step descriptions to produce realistic, varied, and physically plausible assembly or human motions; includes synthesising large-scale, diverse trajectories beyond existing motion-capture distributions using geometric, kinematic, and physical priors.

proceduralanimation

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Human Motion Video Generation: A Survey

Sep 04, 2025
HX
Haiwei Xue
🏛️ Tsinghua University | 01.AI | Xi’an Jiaotong University | University of Chinese Academy of Sciences | Nanjing University of Science and Technology | Shenzhen Institutes of Advanced Technology | Huawei Noah’s Ark Lab | Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ) | Artificial Intelligence Innovation and Incubation (AI) Institute of Fudan University | Shenzhen University | Carleton University

Existing surveys predominantly focus on isolated methodologies, lacking a systematic examination of the end-to-end pipeline for human motion video generation. To address this gap, we propose the first unified framework encompassing five core stages: input specification, motion planning, generation, optimization, and output rendering—supporting over ten subtasks driven by visual, textual, and audio modalities. Our work systematically reviews 200+ papers and constructs the field’s first comprehensive technical taxonomy. We innovatively investigate the potential of large language models (LLMs) for motion semantic modeling and cross-modal alignment. Furthermore, we integrate state-of-the-art techniques—including diffusion models, generative adversarial networks (GANs), and multimodal fusion—to identify key breakthroughs and release an open-source model library. This study fills a critical void in holistic, cross-cutting research on human motion video generation, providing both theoretical foundations and practical guidelines for applications such as digital avatars.

Addressing gaps across ten sub-tasks and five key phasesExploring large language models' potential for motion generationLack of comprehensive overview in human motion video generation

Must-Read Papers

Most classic and influential ideas
View more

Human Motion Prediction, Reconstruction, and Generation

Feb 21, 2025
CG
Canxuan Gang
🏛️ AI Geeks

This paper presents a systematic review of recent advances in human motion prediction, reconstruction, and generation. Addressing key challenges—including instability in long-horizon prediction, limited reconstruction accuracy, and insufficient physical plausibility and diversity in motion generation—we propose a unified “prediction–reconstruction–generation” co-evolutionary framework. Our method integrates diffusion models with physics-informed dynamical constraints in the loss function to enhance motion realism and biomechanical consistency. Furthermore, we introduce multimodal alignment and fine-grained contextual modeling to improve text-to-motion generation and human-object interaction synthesis. Extensive experiments demonstrate significant improvements over state-of-the-art methods: a 23% reduction in average prediction error for long-horizon motion forecasting, an 18% decrease in MPJPE for 3D pose reconstruction, and a 31% reduction in FID score for motion generation. The framework supports applications in digital avatars, embodied AI, and real-time AR interaction.

Forecasting future human poses from historical dataRecovering 3D human movements from visual inputsSynthesizing realistic motions from textual descriptions

Generating Continual Human Motion in Diverse 3D Scenes

Apr 04, 2023
AM
Aymen Mir
🏛️ University of Tübingen | Max Planck Institute for Informatics | Meta AI Research | University of California, Berkeley

Generating long-horizon, multi-action coherent human motions in 3D scenes remains challenging due to drift accumulation, action discontinuity, and poor scene adaptability. Method: We propose an animator-guided, scene-agnostic iterative generation framework. It establishes a target-centric canonical coordinate system to decouple path planning from motion transition, and employs motion decomposition modeling with coordinate-system reparameterization—enabling zero-shot deployment on pure motion-capture data without scene-aware annotations or fine-tuning. Contribution/Results: To our knowledge, this is the first method to generate drift-free, chained multi-action sequences (e.g., “grasp → sit → lean”) in diverse real-world scanned environments—including HPS, Replica, Matterport, and ScanNet—using only sparse keypoint constraints and a seed motion. Unlike existing 3D navigation approaches, ours requires no scene rendering, geometric encoding, or environment-specific training, achieving superior generalization, motion plausibility, and scene compatibility.

3D Human Motion SimulationAdaptability to Different 3D EnvironmentsCoherence and Naturalness

This study addresses the limitations of traditional physics simulators in robotics—such as restricted expressiveness due to simplifying assumptions, high data costs, and difficulties in modeling complex physical interactions—by systematically reviewing video generation models as embodied world models. Integrating high-fidelity, multimodal-conditioned video synthesis with imitation learning, reinforcement learning, and visual planning frameworks, this work provides the first comprehensive analysis of their potential and limitations in tasks including action prediction, dynamics modeling, and policy evaluation. The review highlights breakthroughs in high-fidelity modeling of physical interactions while identifying key challenges in instruction following, physical consistency, and safety. These insights lay a theoretical foundation and outline future directions for replacing conventional simulators and enabling deployment in safety-critical scenarios.

hallucinationphysics violationrobotics

HUMOS: Human Motion Model Conditioned on Body Shape

Sep 05, 2024
ST
Shashank Tripathi
🏛️ Max Planck Institute for Intelligent Systems | Epic Games

Existing motion generation models often neglect individual body shape variations, relying instead on a generic average human template—leading to physically implausible and anatomically homogeneous motions. To address this, we propose the first generative motion model conditioned on 3D body shape (parameterized by SMPL-X), enabling body-aware motion synthesis without requiring paired motion-capture data. Our approach jointly models the coupling between body shape and motion dynamics, incorporating a cycle-consistency loss, physics-based constraints derived from kinematics and dynamics, and stability regularization. Quantitative evaluation demonstrates consistent superiority over state-of-the-art methods across standard metrics—including FID, Jitter, and Diversity. Qualitative analysis further confirms anatomical plausibility, inertial consistency, and strong cross-body-shape generalization. Overall, our method significantly enhances both the physical realism and inter-individual diversity of synthesized human motion.

Generating realistic human motion for diverse body shapesOvercoming uniform motion limitations in existing modelsTraining motion models with unpaired data using constraints

Latest Papers

What's happening recently
View more

Existing approaches to human motion generation often suffer from physically implausible results due to contact modeling limited to the hands. This work proposes a physics-aware framework that explicitly models the full spectrum of contacts—including interactions between the body and objects, scenes, and self-limb collisions—through a continuous distance-driven force model. By integrating soft physical constraints with force and torque balance mechanisms, the method synthesizes multi-body dynamics-consistent motions. It supports arbitrary surface interactions with both static environments and dynamic objects, significantly enhancing the physical plausibility of generated actions. The approach demonstrates strong generalization in complex, dynamic scenarios and establishes a new benchmark for physically consistent human motion generation.

contact modelingdynamic scenehuman motion synthesis

Current image-to-video generation models struggle to accurately simulate mechanical motion governed by kinematic and geometric constraints, often exhibiting inconsistencies in rigidity preservation, component contact, and motion transmission. This work proposes MechVerse—the first benchmark dataset specifically designed for mechanical assembly scenarios—which systematically defines and quantifies mechanical motion consistency in video generation. The benchmark encompasses three levels of mechanism complexity and establishes a multi-tiered evaluation framework integrating synthetic data, structured prompts, standard video metrics, instruction-following scores, and human assessments of motion correctness. Experiments reveal that while state-of-the-art models maintain visual fidelity and temporal smoothness, they perform poorly in terms of mechanical plausibility, with error rates rising significantly as coupling complexity increases.

kinematic constraintsmechanical assembliesmotion correctness

Existing motion capture datasets suffer from limited diversity, which constrains the generalization capabilities of generative models on rare, highly dynamic, and compositionally complex actions. To address this limitation, this work proposes a method that leverages large-scale synthetic human motion data combined with physics-based plausibility constraints to jointly expand both the training distribution and the size of the discrete codebook. By reconstructing the VQ-VAE motion tokenizer beyond the confines of real-data distributions, the approach substantially broadens the coverage and compositional capacity of the discrete motion representation space. This leads to consistent performance gains in text-to-motion generation and motion in-betweening tasks, and the enhanced representations can be seamlessly integrated into existing frameworks such as MotionGPT, demonstrating both the effectiveness and generalizability of the proposed representation expansion.

long-tail motionsmotion capturemotion generation

This work addresses the challenge that existing methods struggle to accurately interpret diverse motion categories in textual prompts, limiting the quality of multi-instance video synthesis. The authors propose a training-free motion decomposition framework that disentangles complex motion into three canonical types: static, rigid, and non-rigid. Adopting a “plan-then-generate” paradigm, the approach first infers instance-level shape and positional dynamics through a motion graph during the planning phase, then modulates each motion type in a disentangled manner during generation. This method achieves, for the first time, training-free disentanglement of motion categories, introduces motion graph–guided structured semantic representations, and incorporates model-agnostic modules compatible with various diffusion architectures. Experiments demonstrate significant improvements in motion synthesis quality on real-world benchmarks, effectively enabling compositional video generation with multiple instances, appearances, and motion types.

compositional video generationmotion categoriesmotion factorization

Kinematify: Open-Vocabulary Synthesis of High-DoF Articulated Objects

Nov 03, 2025
JW
Jiawei Wang
🏛️ UC San Diego | Deemos Technology Co., Ltd. | ShanghaiTech University

Existing modeling approaches for high-degree-of-freedom articulated objects (e.g., robots) rely heavily on manual annotations or motion sequences, suffering from poor scalability and labor-intensive data curation. Method: This paper introduces the first end-to-end, open-vocabulary (RGB image or text prompt) automatic modeling framework. It jointly performs topology inference via Monte Carlo Tree Search (MCTS) and geometry-driven optimization for joint parameter estimation—requiring neither motion data nor hand-crafted datasets. Contribution/Results: Our method is the first to synthesize physically consistent and functionally plausible articulated models directly from a single RGB image or natural language description. By decoupling structural inference from parametric estimation, it ensures both topological correctness and kinematic plausibility. Evaluated on synthetic and real-world benchmarks, it achieves significant improvements in registration accuracy (+12.3%) and topology recognition accuracy (+18.7%), demonstrating strong generalization and practical utility.

Automating articulated object synthesis from images or text promptsEstimating joint parameters from static geometric informationInferring kinematic topologies for high-DoF complex objects

Hot Scholars

HQ

Huamin Qu

Chair Professor, Hong Kong University of Science and Technology
Data visualizationHuman-Computer InteractionExplainable AIE-Learning
GL

Guosheng Lin

Nanyang Technological University
Computer VisionMachine Learning
LB

Liefeng Bo

Head of Applied Computer Vision Lab at Alibaba Group
Machine LearningComputer VisionRobotics
DM

Dechao Meng

PhD candidate, Institute of Computing Technology, Chinese Academy of Science
deep learningcomputer vision