๐ค AI Summary
This work addresses the challenge of balancing temporal consistency and detail preservation in Stable Diffusionโbased video generation. The authors propose a motion-adaptive temporal attention mechanism that, without modifying the original frozen model, dynamically adjusts the attention receptive field based on motion estimation. This approach integrates a cascaded UNet injection strategy, temporally correlated noise initialization, and a motion-aware gating mechanism, introducing only 25.8 million (2.9%) additional trainable parameters. Experiments on the WebVid validation set demonstrate competitive generation quality, marking the first demonstration of high-fidelity video synthesis without explicit temporal loss terms. Furthermore, the method reveals a controllable trade-off between noise correlation and motion magnitude, offering new insights into motion-guided generative modeling.
๐ Abstract
We present a motion-adaptive temporal attention mechanism for parameter-efficient video generation built upon frozen Stable Diffusion models. Rather than treating all video content uniformly, our method dynamically adjusts temporal attention receptive fields based on estimated motion content: high-motion sequences attend locally across frames to preserve rapidly changing details, while low-motion sequences attend globally to enforce scene consistency. We inject lightweight temporal attention modules into all UNet transformer blocks via a cascaded strategy -- global attention in down-sampling and middle blocks for semantic stabilization, motion-adaptive attention in up-sampling blocks for fine-grained refinement. Combined with temporally correlated noise initialization and motion-aware gating, the system adds only 25.8M trainable parameters (2.9\% of the base UNet) while achieving competitive results on WebVid validation when trained on 100K videos. We demonstrate that the standard denoising objective alone provides sufficient implicit temporal regularization, outperforming approaches that add explicit temporal consistency losses. Our ablation studies reveal a clear trade-off between noise correlation and motion amplitude, providing a practical inference-time control for diverse generation behaviors.