hierarchical motion modeling

Designs and implements hierarchical representations and models of articulated motion and kinematic chains that decompose movement into layered components (e.g., global/root motion, limb- and joint-level motions, and residual corrections) across spatial and temporal scales. Builds and analyzes kinematic mappings and simulators—forward-kinematics solvers, multibody/kinematic-skeleton parameterizations, and residual correction modules—to synthesize, predict, or evaluate poses, keyframes, and biomechanically constrained motion sequences.

hierarchicalmotionmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Articulated Kinematics Distillation from Video Diffusion Models

Apr 01, 2025
XL
Xuan Li
🏛️ UCLA | NVIDIA

This work addresses three key challenges in text-driven 4D character animation generation: low motion quality, structural inconsistency, and physically implausible dynamics. To this end, we propose a joint-level motion distillation framework. Methodologically, we distill 3D kinematic priors from a pre-trained video diffusion model and integrate them with skeleton-driven representation and Score Distillation Sampling (SDS), enabling joint modeling of expressive motion and skeletal topology consistency. The resulting joint trajectories are inherently compatible with physics-based simulation, circumventing the intrinsic shape instability issues inherent to 4D neural deformation fields. Experiments demonstrate that our approach significantly improves both 3D structural consistency and motion fidelity in text-to-4D generation, achieving state-of-the-art performance across quantitative and qualitative evaluations.

Ensures structural integrity and physical plausibility in articulated motionsGenerates high-fidelity character animations using skeleton-based and generative modelsReduces motion complexity by focusing on joint-level control

Kinematify: Open-Vocabulary Synthesis of High-DoF Articulated Objects

Nov 03, 2025
JW
Jiawei Wang
🏛️ UC San Diego | Deemos Technology Co., Ltd. | ShanghaiTech University

Existing modeling approaches for high-degree-of-freedom articulated objects (e.g., robots) rely heavily on manual annotations or motion sequences, suffering from poor scalability and labor-intensive data curation. Method: This paper introduces the first end-to-end, open-vocabulary (RGB image or text prompt) automatic modeling framework. It jointly performs topology inference via Monte Carlo Tree Search (MCTS) and geometry-driven optimization for joint parameter estimation—requiring neither motion data nor hand-crafted datasets. Contribution/Results: Our method is the first to synthesize physically consistent and functionally plausible articulated models directly from a single RGB image or natural language description. By decoupling structural inference from parametric estimation, it ensures both topological correctness and kinematic plausibility. Evaluated on synthetic and real-world benchmarks, it achieves significant improvements in registration accuracy (+12.3%) and topology recognition accuracy (+18.7%), demonstrating strong generalization and practical utility.

Automating articulated object synthesis from images or text promptsEstimating joint parameters from static geometric informationInferring kinematic topologies for high-DoF complex objects

KinMo: Kinematic-aware Human Motion Understanding and Generation

Nov 23, 2024
PZ
Pengfei Zhang
🏛️ University of California, Irvine | University of Rochester | Imperial College | FlawlessAI

Existing text-driven human motion generation methods rely on global action descriptors (e.g., “running”), failing to capture velocity variations, joint poses, and kinematic–dynamic constraints—leading to semantic ambiguity between text and motion modalities and lacking fine-grained controllability. To address this, we propose a kinematics-aware joint-group decomposition representation and a hierarchical semantic alignment framework. Specifically, we introduce biomechanically constrained joint grouping for the first time; construct the first automatically generated fine-grained text–motion paired dataset; and design a coarse-to-fine hierarchical semantic fusion and generation architecture that enables joint-level interactive encoding and cross-modal alignment. Experiments demonstrate significant improvements in text-to-motion retrieval accuracy—particularly in joint-spatial understanding—and enable high-fidelity, editable local joint motion generation and manipulation.

Addresses modality gap in human motion synthesisEnhances motion generation via hierarchical text-motion alignmentImproves motion understanding with fine-grained descriptions

This work addresses the challenge of achieving hierarchical action representations in humanoid motor control that jointly support structured abstraction and fine-grained manipulation. The authors propose MotionPyramid, which introduces the concept of hierarchical perceptual representations into motor control by recursively stacking latent decoders to learn multiscale action hierarchies directly from motion data: high-level latent variables generate temporally extended action segments, while low-level variables produce immediate whole-body commands. After pretraining, this hierarchy is frozen and employed as a multi-resolution action interface for reinforcement learning policies, augmented with a residual mechanism that fuses coarse-grained action programs with fine-grained corrective commands. Experiments demonstrate that this approach accelerates early-stage learning, enhances movement regularity, and preserves feedback control precision, thereby unifying structured abstraction with high-fidelity motor execution.

action interfaceshierarchical motion representationhumanoid control

Existing whole-body musculoskeletal models suffer from limited muscle counts (<100) and joint degrees of freedom, hindering high-fidelity, real-time coordinated control of 600+ muscles. To address this, we propose MS-HUMAN-700—the first full-body musculoskeletal model featuring 700 anatomically grounded muscle units, 90 body segments, and 206 kinematic joints. We design a hierarchical low-dimensional representation framework that maps the high-dimensional muscle activation space into a learnable latent space. Integrating biomechanical modeling with hierarchical deep reinforcement learning, our method enables closed-loop, muscle-level motor control. In simulation, MS-HUMAN-700 accurately reproduces human gait patterns with state-of-the-art control fidelity. Both the model and algorithm are fully open-sourced, establishing a scalable neuro-muscular control foundation for embodied intelligence and human–machine interaction.

Human Body ModelingMotion ControlMuscle Simulation

Latest Papers

What's happening recently
View more

Existing methods struggle to uniformly model joint motion across arbitrary skeleton topologies, often constrained by fixed architectures or suffering semantic distortion during cross-skeleton transfer. This work proposes a skeleton-aware motion representation that jointly encodes joint motion, bone connectivity, and semantic joint names via a graph Transformer. It introduces a functional joint group correspondence mechanism, a topology-agnostic attention supervision loss, and a joint name dropout strategy to achieve cross-skeleton action semantic alignment. By integrating cross-attention pooling, residual vector quantization, and a MaskGIT generative model, the approach constructs a part-level discrete motion codebook, enabling high-quality motion transfer, text-guided generation, and fine-grained editing. On cross-topology reconstruction tasks, it achieves a normalized MPJPE of 2.75×10⁻², reducing error by 5.8× compared to the strongest baseline.

articulated objectscross-topology representationmotion generation

Simulating large-scale articulated rigid-body systems remains challenging for conventional rigid-body solvers due to geometric nonlinearities and numerical stiffness. This work proposes a co-rotational framework based on Affine Body Dynamics (ABD) that decouples geometric nonlinearities through a linear kinematic mapping and projects high-dimensional body coordinates onto a dual space spanned by the minimal joint degrees of freedom. By combining implicit integration with KKT system solves, the method enforces exact constraint satisfaction and ensures physically accurate motion propagation. It supports diverse topologies—including chains, trees, closed loops, and irregular networks—and leverages pre-factorization of constant-coefficient matrices to achieve significant computational efficiency. The approach enables interactive simulation of systems comprising hundreds of thousands of rigid bodies on a single CPU core, maintaining high stability and accuracy even with large time steps.

articulated assembliesgeometric complexitymulti-body dynamics

This work addresses the gap between geometric 3D human pose estimation and the biomechanical attributes required in rehabilitation and sports science. We propose BioModule, a lightweight, pose-estimator-agnostic temporal Transformer module that can be appended to any existing 3D pose estimator to predict biomechanically meaningful quantities from standard 17-joint skeletons. To enable frame-level cross-modal supervision, we construct the first large-scale aligned dataset and systematically analyze the impact of upstream pose accuracy on downstream biomechanical prediction performance. By integrating anatomical coordinate alignment with the Human3.6M family of datasets, BioModule demonstrates consistent effectiveness across seven state-of-the-art pose estimators, enabling, for the first time, non-invasive and physically interpretable visual biomechanical analysis.

3D human pose estimationbiomechanical attributeshuman motion analysis

This work addresses the challenge of high-quality articulated part reconstruction, segmentation, and kinematic analysis under conditions where the number of parts is unknown a priori and object visibility is limited. The authors propose a dynamic-static decoupling framework that requires no structural priors, leveraging user interaction videos together with an initial static scan to automatically infer the number of parts and assign joints. By introducing a dual-Gaussian scene representation, the method enables high-fidelity rendering and motion-aware segmentation. Furthermore, it integrates motion cues, sequential RANSAC clustering, and kinematic estimation to achieve end-to-end part parsing. Experiments demonstrate that the approach significantly outperforms existing methods on both simple and complex objects, exhibiting strong generalization and robustness.

3D reconstructionarticulated objectsmotion analysis

Existing simulation and embodied intelligence systems struggle to model fine-grained dynamical effects of articulated objects, such as frictional holding, positional sticking, and damped closure. This work proposes a structured three-channel field representation that explicitly captures conservative forces, dry friction, and damping along joint degrees of freedom, enabling inference and composition of interpretable dynamical primitives from vision-language inputs. For the first time, joint dynamics are formulated as a composable, differentiable function space compatible with physics-based simulation. By integrating shape-constrained piecewise cubic Hermite interpolation (PCHIP) with gradient-based optimization, the method achieves realistic and controllable modeling of complex mechanical behaviors. The framework provides a unified interface for dynamics inference, editing, and optimization, and will be accompanied by open-sourced code and example assets.

articulated objectsdynamical modelingfine-grained dynamics

Hot Scholars

YS

Yanan Sui

Tsinghua University
Optimization and ControlMachine LearningNeural EngineeringRobotics
NN

Nassir Navab

Professor of Computer Science, Technische Universität München
CK

C. Karen Liu

Professor of Computer Science, Stanford University
Computer GraphicsRobotics.
JL

Jiaoyang Li

Assistant Professor at Robotics Institute, Carnegie Mellon University
Artificial IntelligenceMulti-Agent/Robot SystemsHeuristic SearchAutomated Planning
KK

Kento Kawaharazuka

The University of Tokyo
HumanoidBiomimeticsTendon-drivenSoft Robotics