🤖 AI Summary
This study addresses the issue of bidirectional video diffusion models generating content that violates physical laws and logical rules by proposing the S2PD framework. This work introduces a novel serial-parallel hybrid diffusion architecture: an autoregressive strategy is employed during high-noise stages to ensure valid state transitions, while parallel denoising is adopted in low-noise stages to enhance sampling efficiency, thereby jointly preserving temporal causality and global consistency. The method is built upon a pixel-space Diffusion Transformer (DiT), integrating LoRA fine-tuning with causal attention mechanisms for efficient training. Experimental results demonstrate that S2PD significantly outperforms existing baselines in terms of physical logic adherence, temporal stability, and sampling efficiency.
📝 Abstract
Bidirectional video diffusion models denoise entire videos in parallel, yet when trained on effectively unlimited in-distribution data from procedural generators, continue to violate physical laws and simple symbolic rules. We introduce Serial-to-Parallel Diffusion (S2PD), which performs autoregressive diffusion at high noise before switching to parallel diffusion at low noise. The autoregressive phase provides the serial computation needed to coordinate interdependent events and produce valid state transitions while the parallel phase jointly refines the entire video and reduces sampling time relative to fully serial generation. We implement S2PD with two architectures: a pixel-space diffusion transformer trained from scratch and a pretrained video model adapted through LoRA fine-tuning with causal attention. Across games, physical simulations, and real video, S2PD follows rules more reliably than matched bidirectional baselines and generates videos with greater temporal stability and sampling efficiency than other serial methods.