Motion-Adaptive Temporal Attention for Lightweight Video Generation with Stable Diffusion

๐Ÿ“… 2026-03-18
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of balancing temporal consistency and detail preservation in Stable Diffusionโ€“based video generation. The authors propose a motion-adaptive temporal attention mechanism that, without modifying the original frozen model, dynamically adjusts the attention receptive field based on motion estimation. This approach integrates a cascaded UNet injection strategy, temporally correlated noise initialization, and a motion-aware gating mechanism, introducing only 25.8 million (2.9%) additional trainable parameters. Experiments on the WebVid validation set demonstrate competitive generation quality, marking the first demonstration of high-fidelity video synthesis without explicit temporal loss terms. Furthermore, the method reveals a controllable trade-off between noise correlation and motion magnitude, offering new insights into motion-guided generative modeling.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Deep Generative Models & AutoencodersNatural Language Processing: Generation

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
๐Ÿ“ Abstract
We present a motion-adaptive temporal attention mechanism for parameter-efficient video generation built upon frozen Stable Diffusion models. Rather than treating all video content uniformly, our method dynamically adjusts temporal attention receptive fields based on estimated motion content: high-motion sequences attend locally across frames to preserve rapidly changing details, while low-motion sequences attend globally to enforce scene consistency. We inject lightweight temporal attention modules into all UNet transformer blocks via a cascaded strategy -- global attention in down-sampling and middle blocks for semantic stabilization, motion-adaptive attention in up-sampling blocks for fine-grained refinement. Combined with temporally correlated noise initialization and motion-aware gating, the system adds only 25.8M trainable parameters (2.9\% of the base UNet) while achieving competitive results on WebVid validation when trained on 100K videos. We demonstrate that the standard denoising objective alone provides sufficient implicit temporal regularization, outperforming approaches that add explicit temporal consistency losses. Our ablation studies reveal a clear trade-off between noise correlation and motion amplitude, providing a practical inference-time control for diverse generation behaviors.
Problem

Research questions and friction points this paper is trying to address.

video generation
temporal attention
motion adaptation
lightweight model
Stable Diffusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

motion-adaptive temporal attention
lightweight video generation
frozen Stable Diffusion
temporal receptive field
implicit temporal regularization
๐Ÿ”Ž Similar Papers