In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational redundancy in cache reconstruction and long-horizon appearance-motion drift inherent in autoregressive video diffusion. To this end, it proposes FlashForward, a framework that introduces an online KV cache reuse mechanism to directly leverage cached states from the denoising phase, thereby eliminating additional forward-pass overhead. Furthermore, a sparse clean anchor conditioning strategy is incorporated to stabilize generation trajectories, complemented by multi-stage pipeline parallelism and a dual-scale memory architecture to enhance efficiency. Experimental results demonstrate that FlashForward accelerates generation speed by 1.16× to 2.92× across models ranging from 1.3B to 14B parameters while achieving superior VBench scores. Notably, the framework successfully enables the synthesis of high-quality, temporally consistent videos up to 65 seconds in duration.
📝 Abstract
Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less-noisy key--value (KV) cache by additional forwards to build the cache without advancing an output latent. However, every denoising forward itself already computes the in-flight KV of the current chunk. We introduce FlashForward, which directly reuses this cache to avoid the heavy cache-update-only model forwards. After the current chunk completes one denoising stage, its stage-specific cache is already available for the next chunk. Assigning one GPU to each stage therefore lets different chunks occupy different stages concurrently. This early availability has a quality cost: the resulting stage-matched history is noisy, causing appearance and motion drift among chunks. To complement it, FlashForward produces sparse auxiliary clean anchor latents before the corresponding region is generated so the generation trajectories can be stabilized by this two-sided conditioning. The two memories operate at different temporal scales: sparse clean anchor KV supplies coarse, long-range two-sided structural guidance, while dense stage-matched history preserves fine, recent evolution. With up to four GPUs, FlashForward runs $1.16$--$1.69\times$ faster than HiAR and $1.42$--$2.92\times$ faster than Self-Forcing for 16 FPS videos of 20 seconds or longer across 1.3B and 14B backbone scales at 480p and 720p. On VBench, for the 1.3B model at 480p, it achieves higher scores and remains stable at longer durations, demonstrating that FlashForward generates high-quality and temporally consistent videos across durations of 20s, 35s and 65s at a much faster generation speed.
Problem

Research questions and friction points this paper is trying to address.

autoregressive video diffusion
KV cache
generation speed
temporal consistency
appearance and motion drift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Autoregressive Video Diffusion
In-Flight KV Cache
Clean Anchors
Pipeline Parallelism
Dual-Memory Conditioning