Where Compute Matters: Heterogeneous Attention for Efficient Video Diffusion

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational cost of self-attention in video diffusion models and its difficulty in accommodating regional denoising variations by proposing HetA-DiT, a heterogeneous attention mechanism. This method introduces a novel content- and timestep-adaptive heterogeneous routing strategy that employs a lightweight uncertainty branch to predict token difficulty, dynamically assigning global or local attention accordingly. It requires only a single parameter to control the quality-efficiency trade-off and incurs no additional Transformer overhead during inference. Evaluations on the Wan model series demonstrate that applying dense attention to merely 20% of tokens maintains competitive generation quality on benchmarks such as VBench while significantly improving computational efficiency. Furthermore, the proposed approach remains compatible with few-step distillation techniques, offering a practical solution for accelerating video diffusion models without compromising performance.
📝 Abstract
Efficient video generation requires reducing the quadratic cost of self-attention over long spatio-temporal token sequences. Existing efficient-attention methods typically apply the same computation pattern to every token, even though denoising difficulty varies substantially across video regions and evolves throughout the generation process. We introduce HetA-DiT, a heterogeneous attention mechanism that adaptively allocates computation according to token difficulty. A lightweight uncertainty branch predicts a token-wise estimate of denoising difficulty, which is used to route uncertain tokens through dense global attention while processing more reliable tokens with efficient local attention. The resulting routing is content- and timestep-adaptive, retains global context where it matters most, and provides a single parameter for controlling the quality-efficiency trade-off. HetA-DiT is compatible with few-step distribution-matching distillation and introduces no additional Transformer evaluation at inference time by reusing uncertainty estimates from the preceding denoising step. We evaluate the method on DMD-distilled Wan2.2-5B and Wan2.1-1.3B models. Across VBench, VBench-2.0, and human preference evaluation, HetA-DiT maintains competitive generation quality while routing only approximately 20% of tokens through dense attention.
Problem

Research questions and friction points this paper is trying to address.

Video Diffusion
Efficient Attention
Self-Attention
Heterogeneous Computation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Heterogeneous Attention
Video Diffusion
Uncertainty Routing
Efficient Generation
Token Difficulty
🔎 Similar Papers