🤖 AI Summary
This study addresses the high inference latency of full spatiotemporal attention in video diffusion Transformers (DiTs) and the limited flexibility of existing sparse attention methods. We formally characterize, for the first time, the periodic stripe patterns inherent in DiT attention and propose Parameterized Stripe Attention (PSA). This approach unifies the encoding of diverse sparsity patterns and generates efficient masks. By integrating a training-free offline search algorithm with FlashAttention-3-level CUDA kernels, PSA enables efficient processing of all sparsity patterns on a single hardware implementation, thereby overcoming the flexibility-efficiency bottleneck. Evaluated on HunyuanVideo and Wan 2.1, our method achieves 1.57× and 1.37× end-to-end speedups, respectively, with negligible degradation in visual quality.
📝 Abstract
Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits \textbf{periodic diagonal stripe structures} along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present {\bf PSA}, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan~2.1 demonstrate that PSA achieves 1.57$\times$ and 1.37$\times$ end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.