Spatial-Temporal Multi-Scale Quantization for Flexible Motion Generation

📅 2025-08-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing approaches to human motion generation face two key bottlenecks: difficulty in modeling multi-scale motion patterns and limited compositional flexibility of discrete representations. This paper introduces MSQ—the first spatiotemporal multi-scale motion quantization framework—which extracts part-level spatial features via multiple encoders and jointly applies temporal interpolation and vector quantization to compress continuous motion into multi-granularity discrete tokens. Its core innovation lies in enabling zero-shot, fine-tuned-free token stitching and recombination across scales, significantly enhancing generative flexibility and cross-task generalization. MSQ integrates generative masked modeling to establish an end-to-end discrete representation learning paradigm. Extensive experiments demonstrate that MSQ outperforms state-of-the-art methods on diverse benchmarks for motion generation, editing, and conditional control tasks. Both quantitative metrics and qualitative analyses consistently validate its superiority.

Technology Category

Computer Vision: Motion & TrackingMachine Learning: Multimodal LearningNatural Language Processing: Generation

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomySearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
Despite significant advancements in human motion generation, current motion representations, typically formulated as discrete frame sequences, still face two critical limitations: (i) they fail to capture motion from a multi-scale perspective, limiting the capability in complex patterns modeling; (ii) they lack compositional flexibility, which is crucial for model's generalization in diverse generation tasks. To address these challenges, we introduce MSQ, a novel quantization method that compresses the motion sequence into multi-scale discrete tokens across spatial and temporal dimensions. MSQ employs distinct encoders to capture body parts at varying spatial granularities and temporally interpolates the encoded features into multiple scales before quantizing them into discrete tokens. Building on this representation, we establish a generative mask modeling model to effectively support motion editing, motion control, and conditional motion generation. Through quantitative and qualitative analysis, we show that our quantization method enables the seamless composition of motion tokens without requiring specialized design or re-training. Furthermore, extensive evaluations demonstrate that our approach outperforms existing baseline methods on various benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Captures multi-scale motion for complex pattern modeling
Enhances compositional flexibility in motion generation
Improves motion editing, control, and conditional generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-scale spatial-temporal quantization for motion
Distinct encoders for varying spatial granularities
Generative mask modeling for motion editing
🔎 Similar Papers
No similar papers found.
Z
Zan Wang
School of Computer Science & Technology, Beijing Institute of Technology
J
Jingze Zhang
State Key Laboratory of General Artificial Intelligence, BIGAI
Y
Yixin Chen
State Key Laboratory of General Artificial Intelligence, BIGAI
Baoxiong Jia
Baoxiong Jia
Ph.D. in Computer Science, UCLA
Computer VisionArtificial Intelligence
W
Wei Liang
Yangtze Delta Region Academy of Beijing Institute of Technology, Jiaxing
S
Siyuan Huang
State Key Laboratory of General Artificial Intelligence, BIGAI