From Slow Bidirectional to Fast Autoregressive Video Diffusion Models

📅 2024-12-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing bidirectional video diffusion models suffer from prohibitive computational overhead due to global frame attention, hindering real-time deployment in interactive applications such as gaming. To address this, we propose the first causal autoregressive video diffusion Transformer, overcoming the error accumulation bottleneck inherent in autoregressive generation via three key innovations: (1) Distribution-Matching Distillation (DMD) in the video domain, enabling faithful trajectory-level knowledge transfer from a teacher to a student model; (2) ODE-trajectory-guided student initialization, improving the plausibility of initial latent states; and (3) an asymmetric causal teacher–student supervision scheme, balancing generation fidelity and inference efficiency. Our model achieves streaming video synthesis at 9.4 FPS on a single GPU with only four sampling steps—down from 50—while setting a new SOTA score of 84.27 on VBench-Long. It supports zero-shot long-video generation and multimodal streaming tasks including video-to-video and image-to-video translation.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Deep Generative Models & AutoencodersNatural Language Processing: Generation

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGEconomics, Online Markets and Human Computation: Economic ramifications for generative AI infrastructure and applicationsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Current video diffusion models achieve impressive generation quality but struggle in interactive applications due to bidirectional attention dependencies. The generation of a single frame requires the model to process the entire sequence, including the future. We address this limitation by adapting a pretrained bidirectional diffusion transformer to an autoregressive transformer that generates frames on-the-fly. To further reduce latency, we extend distribution matching distillation (DMD) to videos, distilling 50-step diffusion model into a 4-step generator. To enable stable and high-quality distillation, we introduce a student initialization scheme based on teacher's ODE trajectories, as well as an asymmetric distillation strategy that supervises a causal student model with a bidirectional teacher. This approach effectively mitigates error accumulation in autoregressive generation, allowing long-duration video synthesis despite training on short clips. Our model achieves a total score of 84.27 on the VBench-Long benchmark, surpassing all previous video generation models. It enables fast streaming generation of high-quality videos at 9.4 FPS on a single GPU thanks to KV caching. Our approach also enables streaming video-to-video translation, image-to-video, and dynamic prompting in a zero-shot manner. We will release the code based on an open-source model in the future.
Problem

Research questions and friction points this paper is trying to address.

Video Production Models
Bidirectional Attention Mechanism
Real-time Processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distribution Matching Distillation
Asymmetric Teacher-Student Training
KV Cache Technique
🔎 Similar Papers
No similar papers found.