BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the error accumulation and insufficient cross-modal consistency in bidirectional motion-text generation caused by the fixed ordering of autoregressive models. To this end, it proposes a unified masked discrete diffusion framework. Methodologically, this study pioneers a decoupled training strategy for unimodal and cross-modal representations and introduces a generation-aware self-correction mechanism to suppress error propagation during inference. Combined with a two-stage training pipeline and iterative bidirectional prediction, the approach achieves high-quality bidirectional modeling. Experimental results demonstrate that the proposed method attains competitive performance on bidirectional tasks across both the HumanML3D and KIT-ML datasets, effectively validating the superiority of the framework.
📝 Abstract
Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bidirectional dependencies between language and motion, allowing early prediction errors to persist as fixed context and degrade both temporal coherence and cross-modal consistency. Masked discrete diffusion, which models sequences through iterative bidirectional prediction, offers a natural remedy. We therefore propose BiMoGen (Bidirectional Motion-text Generation), a unified masked discrete diffusion framework for bidirectional motion-text modeling. To stabilize training, we design Decoupled Uni- and Cross-Modal Training, in which masked pretraining first establishes cross-modal correspondence on paired motion-text sequences, after which supervised fine-tuning specializes the model for bidirectional generation. Masked diffusion nonetheless introduces its own source of error, as the model is trained on clean ground-truth context yet encounters self-generated and potentially erroneous context at inference, with errors committed under heavily masked states propagating through subsequent steps. We further introduce Generation-Aware Self-Correction that exposes the model to its own predictions during training and applies correction passes at early sampling steps to revise unreliably committed tokens. Extensive experiments on HumanML3D and KIT-ML demonstrate competitive performance on both tasks, validating the effectiveness of the proposed two-stage training and self-correction designs. The project page is available at https://wengwanjiang.github.io/BiMoGen-Page.
Problem

Research questions and friction points this paper is trying to address.

text-to-motion generation
motion-to-text captioning
autoregressive modeling
masked discrete diffusion
error propagation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked Discrete Diffusion
Bidirectional Motion-Text Generation
Decoupled Training
Generation-Aware Self-Correction
Unified Framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wanjiang Weng
Department of Computer Science and Engineering, Southeast University, Nanjing, China
Yongliang Wu
Yongliang Wu
Southeast University
Vision-Language Model
Xiaofeng Tan
Xiaofeng Tan
Research Intern at Tencent; Master at Southeast University; Dual BSc at Shenzhen Unversity.
AIGCRLHF
Xingyu Zhu
Xingyu Zhu
Princeton University
W
Wenbo Zhu
Opus AI
H
Hongsong Wang
Department of Computer Science and Engineering, Southeast University, Nanjing, China