Training-Free Sparse Attention for Fast Video Generation via Offline Layer-Wise Sparsity Profiling and Online Bidirectional Co-Clustering

πŸ“… 2026-03-19
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing training-free sparse attention methods for video generation overlook inter-layer heterogeneity and query-key coupling, struggling to balance generation quality and inference efficiency. This work proposes the SVOO framework, which reveals for the first time that attention sparsity is an intrinsic property varying across layers. Building on this insight, SVOO introduces a two-stage, training-free sparsification mechanism: an offline per-layer sensitivity analysis determines each layer’s inherent pruning ratio, followed by an online bidirectional collaborative clustering strategy that enables block-level sparse attention. Without any additional training, the method achieves up to 1.93Γ— acceleration across seven mainstream video generation models while maintaining a PSNR of 29 dB on the Wan2.1 dataset, significantly outperforming existing approaches.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Mixture of Experts (MoE)Natural Language Processing: Generation

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Vertical and domain-specific searchGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphs
πŸ“ Abstract
Diffusion Transformers (DiTs) achieve strong video generation quality but suffer from high inference cost due to dense 3D attention, leading to the development of sparse attention technologies to improve efficiency. However, existing training-free sparse attention methods in video generation still face two unresolved limitations: ignoring layer heterogeneity in attention pruning and ignoring query-key coupling in block partitioning, which hinder a better quality-speedup trade-off. In this work, we uncover a critical insight that the attention sparsity of each layer is its intrinsic property, with minor effects across different inputs. Motivated by this, we propose SVOO, a training-free Sparse attention framework for fast Video generation via Offline layer-wise sparsity profiling and Online bidirectional co-clustering. Specifically, SVOO adopts a two-stage paradigm: (i) offline layer-wise sensitivity profiling to derive intrinsic per-layer pruning levels, and (ii) online block-wise sparse attention via a novel bidirectional co-clustering algorithm. Extensive experiments on seven widely used video generation models demonstrate that SVOO achieves a superior quality-speedup trade-off over state-of-the-art methods, delivering up to $1.93\times$ speedup while maintaining a PSNR of up to 29 dB on Wan2.1.
Problem

Research questions and friction points this paper is trying to address.

sparse attention
video generation
layer heterogeneity
query-key coupling
diffusion transformers
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free
sparse attention
layer-wise sparsity profiling
bidirectional co-clustering
video generation
πŸ”Ž Similar Papers
No similar papers found.