🤖 AI Summary
Existing video generation methods struggle to maintain physical consistency and causal plausibility over long temporal sequences. This work proposes a training-free framework that, for the first time, integrates explicit physical reasoning into the generation process through structured intermediate representations—specifically, causally consistent keyframes and object-centric motion trajectories. These representations are derived by leveraging a vision-language model to parse textual prompts and are subsequently employed as soft constraints to guide the inference of a pretrained video diffusion model. Without requiring any additional training, the approach substantially enhances the physical plausibility and temporal coherence of generated videos in dynamics-intensive scenarios while preserving high perceptual quality.
📝 Abstract
Recent advances in diffusion-based video generation have significantly improved visual quality and short-term temporal coherence. However, existing methods still struggle to produce videos with physically consistent and causally plausible dynamics, especially in scenarios involving long-horizon interactions. This limitation arises from the fact that video diffusion models primarily learn physical consistency implicitly, while vision-language models can directly model physical laws. Based on this idea, in this work, we propose \textbf{CausalMotion}, a training-free framework that injects explicit physical reasoning into video generation through structured intermediate representations. Our key idea is to decouple reasoning from generation by leveraging a vision-language model to decompose a text prompt into a sequence of causally consistent keyframes and object-centric motion trajectories. These representations are then aligned and integrated as soft constraints to guide a pretrained video diffusion model during inference. This design enables explicit modeling of object dynamics and causal transitions without requiring additional training or supervision. Extensive experiments show that our method consistently improves physical plausibility and temporal coherence, particularly in dynamics-intensive scenarios, while maintaining high perceptual video quality.