🤖 AI Summary
This study addresses the quality degradation and efficiency bottlenecks in diffusion Transformer-based video generation caused by approximation errors in sparse attention. To this end, we propose a training-free acceleration framework that introduces a novel quadratic importance estimation based on query-key dot-product interactions, coupled with a signal-to-noise ratio (SNR)-guided dynamic sparsity scheduling strategy. Furthermore, an irregular cluster-tail merging mechanism is designed to optimize GPU kernel utilization. Evaluated on HunyuanVideo and Wan2.1, the proposed framework achieves 2.03× and 1.55× speedups, respectively, while preserving high-fidelity generation quality with PSNR scores of 31.017 dB and 29.301 dB. These results demonstrate that our approach effectively realizes the synergistic optimization of inference efficiency and generation quality.
📝 Abstract
Training-free sparse attention offers a practical acceleration solution to Diffusion Transformers (DiTs) via reducing computations without fine-tuning. It typically involves estimating the importance of query-key regions and deriving sparse masks to compute only the important candidates, which inevitably introduces approximation errors that may degrade generation quality. To better balance the efficiency-quality trade-off, we propose iCATS, integrating improved importance estimation and sparse mask construction with an efficient hardware execution strategy. Specifically, for importance estimation, unlike previous works that perform independent clustering over query and key tokens based on feature similarity to estimate attention scores, iCATS demonstrates that clustering based on query-key dot-product interactions is more accurate and further reformulates this objective as a simple quadratic form for low-cost computation. For sparse mask construction, instead of using a fixed top-p rule, we observe that tolerance to sparse approximation errors varies across denoising timesteps and therefore introduce an SNR-guided sparsity schedule to adjust sparsity dynamically, leading to higher accuracy. Finally, for hardware execution, we devise a tail-merging strategy to reduce padding overhead caused by irregular cluster sizes, improving GPU kernel utilization. Extensive experiments show that iCATS achieves $2.03\times$ acceleration with 31.017 dB PSNR on HunyuanVideo-T2V-13B and $1.55\times$ acceleration with 29.301 dB PSNR on Wan2.1-T2V-14B, delivering a state-of-the-art efficiency-quality trade-off.