Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational overhead and learning signal dilution caused by redundant tokens in full-attention mechanisms for high-resolution joint audio-visual generation. To overcome these limitations, this work proposes a dynamic sparse attention framework that adaptively allocates non-uniform attention block shapes through spatiotemporal macro-partitioning and cross-modal feature analysis, guided by visual variance and audio-visual coupling strength to achieve semantically coherent sparsification. Compared with full-attention baselines, the proposed method yields a 2.5× training acceleration while significantly enhancing generation quality, thereby establishing a new paradigm for efficient multimodal video synthesis.
📝 Abstract
Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, diluting learning signals for informative content and disrupting pretrained priors. Existing sparse attention methods either target training-free acceleration or overlook the unique structure of joint video-audio data, where cross-modal interactions are inherently concentrated around sound-producing regions. To address this, we propose Prism, a dynamic sparse attention framework for natively training joint video-audio generation models at 2K. In particular, Prism organizes the token sequence into spatiotemporal macro-zones, enabling the attention structure to adapt to local content. For each zone, it estimates local information structure via video feature variance along the channel and feature norms from the audio-to-video cross-attention, jointly capturing how visual content varies directionally and how strongly audio influences each visual region. Based on these signals, Prism dynamically assigns a tailored block shape to each zone, applying finer partitioning along axes of rapid visual content variation and strong audio-visual coupling. This encourages tokens within each block to remain semantically coherent, allowing block-level features to capture both visual content and joint video-audio interaction patterns. Prism further adopts a hybrid block selection strategy to dynamically determine per-query sparsity. Experiments show that Prism achieves 2.5$\times$ training speedup compared to full attention, while surpassing it in generation quality.
Problem

Research questions and friction points this paper is trying to address.

joint video-audio generation
sparse attention
high-resolution training
computational efficiency
cross-modal interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Sparse Attention
Joint Video-Audio Generation
Spatiotemporal Macro-zones
Hybrid Block Selection
Native 2K Training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.