Parameterized Stripe Attention for Efficient Video Generation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference latency of full spatiotemporal attention in video diffusion Transformers (DiTs) and the limited flexibility of existing sparse attention methods. We formally characterize, for the first time, the periodic stripe patterns inherent in DiT attention and propose Parameterized Stripe Attention (PSA). This approach unifies the encoding of diverse sparsity patterns and generates efficient masks. By integrating a training-free offline search algorithm with FlashAttention-3-level CUDA kernels, PSA enables efficient processing of all sparsity patterns on a single hardware implementation, thereby overcoming the flexibility-efficiency bottleneck. Evaluated on HunyuanVideo and Wan 2.1, our method achieves 1.57× and 1.37× end-to-end speedups, respectively, with negligible degradation in visual quality.
📝 Abstract
Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits \textbf{periodic diagonal stripe structures} along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present {\bf PSA}, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan~2.1 demonstrate that PSA achieves 1.57$\times$ and 1.37$\times$ end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.
Problem

Research questions and friction points this paper is trying to address.

Video Generation
Diffusion Transformers
Sparse Attention
Inference Latency
Spatio-temporal Attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parameterized Stripe Attention
Diffusion Transformers
Sparse Attention
CUDA Kernel Optimization
Training-free Search
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xingyu Jia
Alibaba Group
Baole Ai
Baole Ai
Alibaba
Ang Wang
Ang Wang
Alibaba
K
Kang Zhao
Alibaba Group
Y
Yong Li
Alibaba Group