Score
Designs and implements systems that generate temporal motion sequences from natural-language inputs, i.e., models that map text or language embeddings to parameterized motion trajectories such as fixed-length 3D skeleton or gait sequences. Work includes building text encoders and motion decoders/parameterizations, conditioning and diversity techniques (for example semantic augmentation), and analyses of sample fidelity, temporal coherence, and utility for downstream tasks like data augmentation.
This paper addresses the cross-modal generation problem of text-to-human-motion synthesis, aiming to enable flexible, fine-grained linguistic control over animated avatars. Methodologically, it presents a systematic survey tracing the field’s evolution—from action-label-conditioned prediction to end-to-end language-driven paradigms—and introduces, for the first time, a two-dimensional taxonomy: “model architecture (VAE/diffusion/hybrid) × motion representation (discrete/continuous).” Leveraging benchmarks including KIT-ML and HumanML3D, the work unifies variational autoencoders, diffusion modeling, motion tokenization, and multimodal alignment techniques into a coherent evaluation framework. Key contributions include: (i) clarifying core challenges—semantic alignment, temporal coherence, and motion realism; (ii) establishing empirical performance boundaries; and (iii) advancing methodological standardization and reproducibility. The study provides foundational technical support for applications in VR, gaming, human–computer interaction, and embodied AI.
This work addresses the challenge of achieving precise alignment between motion dynamics and semantic content in text-driven human motion generation. To this end, the authors propose the MLA-Gen framework, which integrates global motion priors with fine-grained local textual conditions to jointly model general motion patterns and detailed semantic correspondence. The study is the first to identify an “attention collapse” phenomenon during generation and introduces SinkRatio—a novel metric to quantitatively assess alignment quality. Building upon this insight, the authors design alignment-aware masking and attention modulation strategies to refine the motion distribution. Extensive experiments demonstrate that MLA-Gen significantly outperforms strong existing baselines across multiple benchmarks, achieving state-of-the-art performance in both motion quality and text-motion semantic alignment.
This work addresses text-driven human motion generation, aiming to enhance precise natural language control over complex, anthropomorphic action sequences. We propose an LLM-empowered semantic alignment paradigm and introduce an end-to-end text-to-motion framework that jointly optimizes text embeddings and joint trajectories via a hybrid architecture integrating autoregressive modeling, diffusion processes, and Transformer-based sequence learning. We establish, for the first time, a comprehensive three-dimensional evaluation framework assessing generation quality, efficiency, and controllability, and systematically chart the technical evolution of the field. Experiments demonstrate significant improvements in motion semantic fidelity and contextual coherence. The approach exhibits strong practical potential in medical rehabilitation, humanoid robotics, and animation production, offering a novel methodological foundation for lightweight and efficient motion generation.
Existing text-to-motion generation methods rely on end-to-end embeddings, struggling to simultaneously achieve temporal precision, fine-grained detail, and interpretability. This work proposes LaMoGen, a novel framework that introduces LabanLite—a lightweight symbolic system derived from Labanotation—and establishes a three-stage “text–symbol–motion” generation pipeline. For the first time, large language models (LLMs) are leveraged at the symbolic level to perform motion reasoning and recombination. This approach enables an interpretable mapping between language and motion, significantly outperforming state-of-the-art methods on both a newly curated Labanotation benchmark and two public datasets. The proposed method achieves substantial improvements in motion interpretability, controllability, and alignment between textual descriptions and generated motions.
MotionGlot addresses the challenges of cross-modal motion generation—namely, inconsistent multi-dimensional action spaces across heterogeneous agents (e.g., quadrupeds and humans), scarcity of high-quality annotated data, and difficulty in text-motion alignment—by adapting large language model (LLM) training paradigms to motion synthesis. To this end, it introduces: (1) a multi-entity motion-space alignment mechanism; (2) a text-motion joint embedding framework with instruction fine-tuning; and (3) the first directionally annotated quadruped locomotion dataset and a large-scale, scenario-aware human motion prompting corpus. Evaluated on six generation tasks, MotionGlot achieves an average 35.3% improvement over prior methods and demonstrates end-to-end deployment on a physical quadruped robot. This work establishes a novel paradigm for text-driven, general-purpose motion generation across diverse embodied agents.
This work addresses the lack of fine-grained, expressive, and interpretable natural language descriptions for 3D human motion. We propose MotionScript—a zero-shot, unsupervised framework that maps motion sequences to structured natural language without training data. Its core is a template-based generation mechanism integrating kinematic priors and semantic rules, enabling systematic production of expressive descriptions covering affective states, stylistic gait, and human–object/human–human interactions. Methodologically, domain-knowledge-guided structured templates collaborate with large language models to jointly optimize linguistic fidelity and motion alignment, substantially improving generalization and diversity in text-driven motion generation. Experiments demonstrate significant performance gains in out-of-distribution motion synthesis. Moreover, we introduce the first large-scale, explainable language annotation resource specifically designed for expressive motion—enabling high-fidelity animation, virtual human simulation, and robot instruction grounding.
Existing video generation models struggle to precisely control the temporal details of complex human motions through text alone, while explicit skeleton-based control requires users to provide lengthy pose sequences, which is labor-intensive and costly. To address this, this work proposes a two-stage cascaded framework: first, an autoregressive text-to-2D-pose model generates motion sequences from textual descriptions; then, a pose-conditioned video diffusion model synthesizes high-quality videos by combining these pose sequences with a reference image. The approach introduces a novel DINO-ALF multi-level reference encoding mechanism to maintain appearance consistency under large pose variations and constructs the first synthetic dataset comprising 2,000 highly controllable acrobatic motion clips. Experiments demonstrate that the proposed method significantly outperforms existing approaches in both pose generation and video synthesis on the newly curated dataset and the Motion-X Fitness benchmark.
Existing approaches typically treat action recognition and text-driven motion generation as separate tasks, overlooking their intrinsic semantic connections. This work proposes CoAMD—a skeleton-coordinate-based autoregressive motion diffusion model—that unifies both tasks within a single framework for the first time. CoAMD leverages a multimodal action recognizer to provide semantic gradient guidance and synthesizes motions through a coarse-to-fine strategy. Furthermore, it establishes a unified training and evaluation framework across tasks using absolute coordinates. Evaluated on 13 benchmarks spanning action recognition, text-to-motion generation, text-motion retrieval, and motion editing, the method achieves state-of-the-art performance across all four tasks, demonstrating its effectiveness and generalizability.
Existing text-to-motion generation methods often rely on a single holistic latent vector, which couples trajectory and joint rotations, limiting multi-task support and leading to error accumulation in long sequences. To address this, this work proposes PRISM, a novel framework that decouples human joints into independent latent tokens, forming a structured spatiotemporal latent space. PRISM introduces a token-level conditioning mechanism with timestep embeddings, enabling unified support for text-driven motion, pose-conditioned generation, and streaming motion synthesis. By integrating a causal variational autoencoder, forward kinematics supervision, and a denoising diffusion strategy, PRISM achieves state-of-the-art performance across multiple benchmarks—including HumanML3D, MotionHub, and BABEL—and demonstrates superior robustness in a 50-scenario user study, significantly mitigating motion drift in long-horizon generation.
This work addresses the limitations of existing large language model (LLM)-based approaches to human motion understanding, which rely on specialized encoders for cross-modal alignment and thereby constrain deep semantic reasoning. The authors propose Structured Motion Description (SMD), a novel framework that, for the first time, translates motion sequences into human-readable natural language text using deterministic biomechanical rules—such as joint angles, body-part movements, and global trajectories. This enables LLMs to directly comprehend and reason about motion semantics without requiring additional encoders. By circumventing conventional cross-modal alignment paradigms, SMD facilitates plug-and-play LLM integration and supports interpretable attention analysis. Experiments demonstrate that SMD achieves state-of-the-art performance, with accuracies of 66.7% and 90.1% on BABEL-QA and HuMMan-QA, respectively, and scores of R@1 = 0.584 and CIDEr = 53.16 on HumanML3D.