MotionMaestro: Masked Tokenization for Unified Motion Generation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the task isolation and absence of a unified multimodal conditioning framework in existing motion generation methods by proposing a unified motion generation framework based on masked tokenization. The framework pioneers the conversion of heterogeneous conditions—such as text, poses, and trajectories—into a unified observation paradigm. By explicitly preserving conditional information through observation graphs and loss mechanisms, it enables fully controllable, multi-task collaborative generation. Furthermore, a three-stage training strategy is designed to optimize the model by integrating a masked motion tokenizer, a latent-space flow matching generator, and conditional constraint techniques. Experimental results demonstrate that the proposed method achieves state-of-the-art performance across multiple motion generation tasks on the RoMo and MotionMillion datasets.
📝 Abstract
Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While existing approaches have achieved remarkable progress, many of them are developed for individual tasks, including text-to-motion, pose-conditioned generation, and trajectory control. Although these tasks involve different types of conditions, a unified framework capable of handling them within a common representation would greatly simplify motion generation systems. We observe that diverse motion conditions can be naturally formulated as different observation patterns over motion sequences, where each task corresponds to a specific masking strategy. Based on this insight, we introduce MotionMaestro, a unified motion generation framework that learns a shared representation for complete motions and heterogeneous partial observations through masked motion tokenization. MotionMaestro employs a three-stage training strategy that first learns a masked motion tokenizer, then refines its reconstruction ability on clean motions, and finally trains a conditional flow-matching generator in the learned latent space. Furthermore, we introduce an observation map and an observation loss to explicitly preserve provided motion conditions during generation. With this unified representation and conditioning mechanism, MotionMaestro supports text-guided and unconditional synthesis, pose conditioning and partial completion, temporal interpolation, trajectory control, and motion continuation. Experiments on the large-scale RoMo and MotionMillion datasets show state-of-the-art performance across diverse motion generation tasks.
Problem

Research questions and friction points this paper is trying to address.

human motion generation
unified framework
multi-task
motion synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked Motion Tokenization
Unified Motion Generation
Conditional Flow-Matching
Observation Map
Three-Stage Training
🔎 Similar Papers
No similar papers found.