MotionSpaceFlow: Representation-Aware Flow Matching in Direct Motion Space

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the quality bottleneck caused by low-dimensional latent space generation and the lack of fine-grained control at the single-frame or joint level in human motion synthesis. To overcome these limitations, this work proposes MSFlow, a framework that eliminates the encoder-decoder paradigm by directly predicting clean motions within the continuous motion space. Specifically, it introduces a representation-aware noise scaling mechanism and an RA-MMDiT architecture that dynamically adapts causal or bidirectional attention flows based on feature types. By integrating flow matching, diffusion transformers, and projection sampling, MSFlow achieves text-driven, high-fidelity human motion generation. Extensive experiments demonstrate that MSFlow attains state-of-the-art performance across multiple datasets while enabling precise zero-shot manipulation of arbitrary joints or frames.
📝 Abstract
Recent advances in diffusion and flow models have substantially improved text-driven human motion generation. Yet most methods generate in low-dimensional, temporally downsampled latent spaces learned primarily for reconstruction, a bottleneck that can limit generation quality and preclude direct manipulation of individual frames and joints. We introduce MotionSpaceFlow (MSFlow), a representation-aware flow-matching framework that predicts clean motion directly in continuous motion space without a learned encoder or decoder. To account for the anisotropic structure of direct motion representations, we propose representation-aware noise scaling and show how the initial Gaussian source scale governs the covariance of intermediate probability-path marginals. We further introduce a Representation-Aware Multimodal Diffusion Transformer (RA-MMDiT), which jointly updates token-level language and full-resolution motion features through joint attention while adapting temporal information flow to the motion representation: causal attention for incremental features defined by frame-to-frame changes, and bidirectional attention for global features such as absolute joint coordinates. Across different datasets and motion representations, MSFlow achieves state-of-the-art text-to-motion performance. Its global representation variant additionally enables zero-shot, inference-time control over any joint or frame through projection sampling without control-conditioned training, delivering leading motion quality with exact constraint satisfaction.
Problem

Research questions and friction points this paper is trying to address.

human motion generation
latent space bottleneck
direct motion manipulation
flow matching
text-to-motion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow Matching
Representation-Aware Noise Scaling
Multimodal Diffusion Transformer
Direct Motion Space
Zero-Shot Control
🔎 Similar Papers
2024-07-11Neural Information Processing SystemsCitations: 0