StreamDiffusion: A Pipeline-level Solution for Real-time Interactive Generation

📅 2023-12-19
🏛️ arXiv.org
📈 Citations: 37
✨ Influential: 4
📄 PDF
🤖 AI Summary
To address the low throughput, high latency, and excessive power consumption of diffusion models in continuous interactive scenarios—such as the metaverse and real-time video streaming—this paper proposes StreamDiffusion, the first streaming diffusion generation framework designed explicitly for real-time interaction. Its core contributions are: (1) Stream Batch, a dynamic input-frame aggregation and decoupling mechanism; (2) Residual Classifier-Free Guidance (RCFG), which reduces guidance overhead while preserving generation quality; and (3) Stochastic Similarity Filtering (SSF), enabling adaptive skipping of redundant computations. The framework integrates parallel input/output queues, seamless Diffusers library compatibility, and fine-grained GPU optimizations. Evaluated on an RTX 4090, StreamDiffusion achieves 91.07 FPS, improves throughput by 59.56×, accelerates inference by 2.05×, and reduces energy consumption by 1.99–2.39×, significantly advancing the deployment of diffusion models in real-time interactive applications.
📝 Abstract
We introduce StreamDiffusion, a real-time diffusion pipeline designed for interactive image generation. Existing diffusion models are adept at creating images from text or image prompts, yet they often fall short in real-time interaction. This limitation becomes particularly evident in scenarios involving continuous input, such as Metaverse, live video streaming, and broadcasting, where high throughput is imperative. To address this, we present a novel approach that transforms the original sequential denoising into the batching denoising process. Stream Batch eliminates the conventional wait-and-interact approach and enables fluid and high throughput streams. To handle the frequency disparity between data input and model throughput, we design a novel input-output queue for parallelizing the streaming process. Moreover, the existing diffusion pipeline uses classifier-free guidance(CFG), which requires additional U-Net computation. To mitigate the redundant computations, we propose a novel residual classifier-free guidance (RCFG) algorithm that reduces the number of negative conditional denoising steps to only one or even zero. Besides, we introduce a stochastic similarity filter(SSF) to optimize power consumption. Our Stream Batch achieves around 1.5x speedup compared to the sequential denoising method at different denoising levels. The proposed RCFG leads to speeds up to 2.05x higher than the conventional CFG. Combining the proposed strategies and existing mature acceleration tools makes the image-to-image generation achieve up-to 91.07fps on one RTX4090, improving the throughputs of AutoPipline developed by Diffusers over 59.56x. Furthermore, our proposed StreamDiffusion also significantly reduces the energy consumption by 2.39x on one RTX3060 and 1.99x on one RTX4090, respectively.
Problem

Research questions and friction points this paper is trying to address.

Enables real-time interactive image generation with high throughput
Reduces redundant computations in diffusion pipeline for faster processing
Optimizes power consumption in continuous input scenarios like Metaverse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transforms sequential denoising into batching process
Introduces residual classifier-free guidance algorithm
Optimizes power with stochastic similarity filter
UC Berkeley | University of Tsukuba | International Christian University | Toyo University | Tokyo Institute of Technology | Tohoku University | MIT
A
Akio Kodaira
UC Berkeley
Chenfeng Xu
Chenfeng Xu
UC Berkeley
Efficient Generative AIEfficient Machine LearningEfficient ComputationAI SystemsRobotics
T
Toshiki Hazama
UC Berkeley
T
Takanori Yoshimoto
University of Tsukuba
K
Kohei Ohno
International Christian University
S
Shogo Mitsuhori
Toyo University
S
Soichi Sugano
Tokyo Institute of Technology
H
Hanying Cho
Tohoku University
Z
Zhijian Liu
MIT
Kurt Keutzer
Kurt Keutzer
Professor of the Graduate School, EECS, University of California, Berkeley
artificial intelligence systemsdeep learningefficient computation