Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of error accumulation and drift in autoregressive video diffusion models during long-sequence generation. To this end, it proposes the SAF training scheme, which for the first time renders historical key/value states differentiable, enabling future losses to directly backpropagate and optimize historical encodings without relying on gradient-free rolling inference. Furthermore, by integrating single-forward parallel denoising with causal masking, the method effectively aligns causal and bidirectional training paradigms. This approach accelerates training by 1.8×, achieving 49.1 FPS on four GPUs, while significantly enhancing both visual quality and temporal consistency in long-horizon video generation.
📝 Abstract
Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visual quality against motion, and that restoring gradients through the history aligns causal training far more closely with bidirectional training. Motivated by these observations, we introduce Self-Aligned Forcing (SAF), a training scheme that aligns the history of each block with the noise level of the block being denoised. Specifically, the history is the K/V produced by preceding blocks at the same denoising stage, so all blocks at a stage can be denoised in a single forward pass under a causal mask. This keeps the noisy history differentiable, allowing future losses to optimize how it is encoded. SAF therefore avoids a separate no-gradient rollout and per-block timestep-zero recaching, training up to 1.8x faster than prior methods with lower memory. At inference, SAF achieves the highest single-GPU throughput among existing methods and keeps one history bank per stage for a multi-GPU pipeline, reaching 49.1 FPS on 4 GPUs. Experiments show superior long-horizon generation with a better balance between visual quality and motion. Project page: https://anonymous.4open.science/w/self-aligned-forcing/.
Problem

Research questions and friction points this paper is trying to address.

autoregressive video diffusion
streaming generation
error accumulation
long-range drift
exposure bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Aligned Forcing
Streaming Video Diffusion
Differentiable Noisy History
Autoregressive Generation
Causal Mask