🤖 AI Summary
This study addresses the high computational cost of self-attention in video diffusion models and its difficulty in accommodating regional denoising variations by proposing HetA-DiT, a heterogeneous attention mechanism. This method introduces a novel content- and timestep-adaptive heterogeneous routing strategy that employs a lightweight uncertainty branch to predict token difficulty, dynamically assigning global or local attention accordingly. It requires only a single parameter to control the quality-efficiency trade-off and incurs no additional Transformer overhead during inference. Evaluations on the Wan model series demonstrate that applying dense attention to merely 20% of tokens maintains competitive generation quality on benchmarks such as VBench while significantly improving computational efficiency. Furthermore, the proposed approach remains compatible with few-step distillation techniques, offering a practical solution for accelerating video diffusion models without compromising performance.
📝 Abstract
Efficient video generation requires reducing the quadratic cost of self-attention over long spatio-temporal token sequences. Existing efficient-attention methods typically apply the same computation pattern to every token, even though denoising difficulty varies substantially across video regions and evolves throughout the generation process. We introduce HetA-DiT, a heterogeneous attention mechanism that adaptively allocates computation according to token difficulty. A lightweight uncertainty branch predicts a token-wise estimate of denoising difficulty, which is used to route uncertain tokens through dense global attention while processing more reliable tokens with efficient local attention. The resulting routing is content- and timestep-adaptive, retains global context where it matters most, and provides a single parameter for controlling the quality-efficiency trade-off. HetA-DiT is compatible with few-step distribution-matching distillation and introduces no additional Transformer evaluation at inference time by reusing uncertainty estimates from the preceding denoising step. We evaluate the method on DMD-distilled Wan2.2-5B and Wan2.1-1.3B models. Across VBench, VBench-2.0, and human preference evaluation, HetA-DiT maintains competitive generation quality while routing only approximately 20% of tokens through dense attention.