SparSP: Exploiting Communication Sparsity for Sequence-Parallel Video DiTs

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations imposed by PCIe bandwidth and inadequate communication optimization for sparse attention during multi-GPU inference of video diffusion models. To overcome these challenges, this work proposes an efficient sparse sequence parallelism system that treats attention sparsity as a communication primitive. The method achieves deep communication-computation synergy through the co-design of dependency-aware token placement, demand-driven KV routing, and a decoupled asynchronous transmission runtime. Experimental results demonstrate that the proposed system yields up to 1.69× end-to-end speedup, reduces communication volume by over 23%, and improves bandwidth utilization by 1.76× compared to NCCL.
📝 Abstract
Diffusion Transformers have become the dominant architecture for video generation. Their substantial computational cost motivates scaling inference across multi-GPU servers, yet efficient scaling remains challenging on commodity GPUs connected via PCIe, whose bandwidth is limited. Although sparse attention substantially reduces computation, its implications for communication remain underexplored. This paper argues that attention sparsity should be treated as a communication primitive. We present SparSP, an efficient sparse sequence parallel communication system that co-designs token distribution, communication routing, and asynchronous execution for sparse video diffusion models. First, Dependency-Aware Placement distributes sequence blocks according to diffusion models'sparse attention patterns. Second, Demand-Directed KV Routing transfers KV blocks directly to requesting GPUs without intermediate relays. Third, a Decoupled Transfer Runtime separates communication from GPU computation to reduce resource contention and maximize effective bandwidth. Our evaluation shows that SparSP improves attention performance by 1.38- 1.5$\times$, achieves an average 1.17$\times$ (up to 1.69$\times$) end-to-end speedup across three representative servers and three video diffusion models, and reduces communication volume by 12.54-23.05%. Moreover, we achieve an average 1.53-1.76$\times$ bandwidth improvement over NCCL.
Problem

Research questions and friction points this paper is trying to address.

Video Diffusion Transformers
Sequence Parallelism
Sparse Attention
Communication Efficiency
Multi-GPU Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Attention
Sequence Parallelism
Diffusion Transformers
Communication Optimization
Video Generation
🔎 Similar Papers
No similar papers found.