Streaming Multi-Track Timeline Control for 3D Human Motion Generation

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of jointly controlling streaming instruction arrival and overlapping action generation in interactive applications by proposing the TimelineControl framework. This work introduces a novel streaming multi-track timeline control paradigm that leverages interval-aware conditioning, causal part-structural representations, and a part-aware denoising diffusion model to enable real-time, precise generation and coordination of concurrent human motions across multiple timeline tracks. To support this approach, we construct the TimelineMotion dataset, which incorporates overlapping instruction intervals. Experimental results demonstrate that the proposed framework significantly outperforms existing baselines in semantic alignment and temporal adherence. Human evaluations further validate the effectiveness of the design, and successful deployment on humanoid robots confirms its practical execution capability.
πŸ“ Abstract
Text-driven human motion generation has advanced substantially, yet most methods assume instructions are available before synthesis. Interactive applications require responding to new instructions while continuing ongoing actions, such as answering a phone while walking. Existing approaches address streaming generation or simultaneous composition without explicitly combining streaming instruction arrival with independently timed, overlapping actions. We introduce streaming multi-track timeline control and propose TimelineControl to incorporate new instructions alongside ongoing actions. Interval-aware conditioning preserves instruction timing, while causal part-structured representations and part-aware denoising coordinate concurrent actions across body regions. We also construct TimelineMotion, a dataset with overlapping instruction intervals and body-part annotations. Experiments on TimelineMotion and MTT demonstrate improved semantic alignment and temporal adherence over evaluated streaming baselines, including models retrained on the same data. Ablations and human evaluations validate our design, complemented by spatial conditioning and humanoid execution demonstrations. Our code, data and models will become publicly available.
Problem

Research questions and friction points this paper is trying to address.

3D human motion generation
streaming instruction
multi-track timeline control
overlapping actions
interactive applications
Innovation

Methods, ideas, or system contributions that make the work stand out.

streaming motion generation
multi-track timeline control
interval-aware conditioning
part-aware denoising
causal part-structured representation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Y
Yangsong Zhang
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
A
Anujith Muraleedharan
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
R
Rikhat Akizhanov
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
G
GΓΌl Varol
LIGM, Γ‰cole des Ponts, IP Paris, Univ Gustave Eiffel, CNRS
Fabio Pizzati
Fabio Pizzati
MBZUAI
computer visiondeep learninggenerative models
Ivan Laptev
Ivan Laptev
Professor at MBZUAI, on leave from INRIA
Computer VisionRoboticsAction RecognitionObject Recognition