iCATS: Fast Video Generation via Interaction-Aware Sparse Attention and Timestep-Adaptive Sparsity
This study addresses the quality degradation and efficiency bottlenecks in diffusion Transformer-based video generation caused by approximation errors in sparse attention. To this end, we propose a training-free acceleration framework that introduces a novel quadratic importance estimation based on query-key dot-product interactions, coupled with a signal-to-noise ratio (SNR)-guided dynamic sparsity scheduling strategy. Furthermore, an irregular cluster-tail merging mechanism is designed to optimize GPU kernel utilization. Evaluated on HunyuanVideo and Wan2.1, the proposed framework achieves 2.03× and 1.55× speedups, respectively, while preserving high-fidelity generation quality with PSNR scores of 31.017 dB and 29.301 dB. These results demonstrate that our approach effectively realizes the synergistic optimization of inference efficiency and generation quality.