PARK: Accurate Block Retrieval for Sparse Attention in Video Diffusion Transformers

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the quality degradation and computational waste caused by block retrieval mismatches in sparse attention for video diffusion models. To this end, we propose PARK, a training-free method that introduces a novel "aggregation-metric" dual-mismatch theory. Specifically, PARK achieves dynamic and precise block retrieval by preserving independent normalization of original queries and employing a QK-aware clustering mechanism that transforms keys based on current queries. Furthermore, it incorporates a custom-designed fused GPU kernel to significantly reduce retrieval overhead. Evaluated on models such as HunyuanVideo, PARK effectively improves retrieval accuracy and accelerates inference while maintaining generation quality, thereby achieving an optimal trade-off between efficiency and performance.
📝 Abstract
Diffusion Transformers (DiTs) have become a dominant architecture for video generation, but their efficiency is limited by the quadratic complexity of full attention. Sparse attention reduces this cost by retrieving important blocks and computing attention only within them, but inaccurate retrieval can either degrade generation quality or yield unnecessary computation. We identify two retrieval mismatches in methods that retrieve blocks using the averaged representations of query and key blocks: (i) query-side aggregation mismatch, where averaging queries before Softmax fails to preserve their individual attention preferences, and (ii) key-side clustering metric mismatch, where standard Euclidean clustering in the original key space can group keys with dissimilar QK scores under the current query, so their average representation may not accurately represent how the current query scores individual keys. These mismatches can lead to inaccurate block retrieval. To address these mismatches, we propose PARK, a training-free sparse attention method for accurate block retrieval. PARK retains every original query, independently normalizes its attention over key blocks, and then averages these distributions within each query block. It also uses information from the current queries to transform keys before clustering, so that keys receiving similar QK scores are grouped together. A fused GPU kernel further reduces the overhead of block retrieval. Experiments on HunyuanVideo and Wan demonstrate that PARK improves block retrieval accuracy and preserves generation quality while accelerating inference, achieving the best quality-efficiency trade-off among the compared sparse attention methods.
Problem

Research questions and friction points this paper is trying to address.

Sparse Attention
Block Retrieval
Video Diffusion Transformers
Retrieval Mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Attention
Block Retrieval
Video Diffusion Transformers
Training-free
Fused GPU Kernel