π€ AI Summary
This study addresses the challenge of jointly controlling streaming instruction arrival and overlapping action generation in interactive applications by proposing the TimelineControl framework. This work introduces a novel streaming multi-track timeline control paradigm that leverages interval-aware conditioning, causal part-structural representations, and a part-aware denoising diffusion model to enable real-time, precise generation and coordination of concurrent human motions across multiple timeline tracks. To support this approach, we construct the TimelineMotion dataset, which incorporates overlapping instruction intervals. Experimental results demonstrate that the proposed framework significantly outperforms existing baselines in semantic alignment and temporal adherence. Human evaluations further validate the effectiveness of the design, and successful deployment on humanoid robots confirms its practical execution capability.
π Abstract
Text-driven human motion generation has advanced substantially, yet most methods assume instructions are available before synthesis. Interactive applications require responding to new instructions while continuing ongoing actions, such as answering a phone while walking. Existing approaches address streaming generation or simultaneous composition without explicitly combining streaming instruction arrival with independently timed, overlapping actions. We introduce streaming multi-track timeline control and propose TimelineControl to incorporate new instructions alongside ongoing actions. Interval-aware conditioning preserves instruction timing, while causal part-structured representations and part-aware denoising coordinate concurrent actions across body regions. We also construct TimelineMotion, a dataset with overlapping instruction intervals and body-part annotations. Experiments on TimelineMotion and MTT demonstrate improved semantic alignment and temporal adherence over evaluated streaming baselines, including models retrained on the same data. Ablations and human evaluations validate our design, complemented by spatial conditioning and humanoid execution demonstrations. Our code, data and models will become publicly available.