🤖 AI Summary
This study addresses the insufficient inference speed of few-step causal video diffusion models in real-time generation by proposing a training-free acceleration framework. Methodologically, it introduces denoising reuse to overcome distillation step limitations, and combines KV-cache window truncation with truncated SVD for attention projections to preserve visual fidelity. These algorithmic advances are integrated with Triton-optimized RoPE and an efficient VAE decoding kernel to achieve system-level speedup. Experimental results demonstrate that the proposed approach attains lossless, extreme acceleration without retraining, reaching 50 FPS on a single H100 GPU and 77 FPS on a GB200 GPU. This work establishes a new throughput standard for real-time video generation.
📝 Abstract
Distilling bidirectional multi-step video diffusion transformers into few-step causal models has become a common approach for streaming video generation. While these few-step students are significantly faster than the teachers they are distilled from, they remain slow for real-time generation. In this work we present UnStep, a training-free wrapper that accelerates few-step causal video models at inference by running them with fewer diffusion transformer (DiT) steps than during distillation and limiting the temporal window retained in the attention KV cache. We propose two inference-only mechanisms to recover quality lost by step reduction and attention windowing: renoising the generated latent frames to a near-clean level and reusing the existing clean-cache pass to refine them, and applying truncated SVD to the DiT attention value and output projections. We also accelerate inference with a quality-preserving runtime stack for the DiT and VAE decoder, including more efficient attention calls and KV indexing, fused Triton RoPE with cached coefficients, and VAE decoding with optimized memory layout, precision, and convolution kernels. By reducing computation and optimizing the runtime stack, UnStep sets a new throughput regime for causal video diffusion, by running substantially faster than current methods, reaching 50 FPS on a single H100 without quality loss, and 77 FPS on GB200, all without retraining.