🤖 AI Summary
This study addresses the limited environmental generalization of existing 3D motion generators and the high computational cost and kinematic instability inherent in two-stage inference pipelines based on video world models. To overcome these limitations, this work constructs a hybrid dataset comprising both synthetic and real-world data and proposes a decoupled noise scheduling strategy that repurposes Cosmos 3 into the first single-stage, scene-aware 3D human motion generation framework. By decoupling the diffusion noise levels for video and motion during model fine-tuning, the approach effectively suppresses temporal jitter. The proposed method significantly enhances text-motion alignment and scene interaction capabilities, achieving a 3.3-fold improvement in inference speed over baseline approaches while maintaining high success rates.
📝 Abstract
We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as Cosmos 3 offer broader environmental priors but are not designed for full-body motion generation; recovering motion from their generated videos requires costly two-stage inference. To address these, we turn Cosmos 3 into a single-stage 3D motion generator. This adaptation has two challenges: the scarcity of paired video--motion data and temporal instability in the generated motion. First, we construct a training dataset combining synthetic video--motion pairs with real videos paired with estimated 3D motion. Second, we propose a shift-decoupled noise schedule that assigns different noise levels to video and motion through shared denoising progress. This design accommodates the different denoising requirements of the two modalities, reducing motion jitter. Experiments on a multi-source interaction benchmark show that World2Motion has better motion--text alignment and scene interaction compared with the evaluated 3D motion generators. It also matches the interaction success rate of the two-stage baseline while achieving approximately 3.3$\times$ faster inference.