Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation

📅 2026-09-17
📈 Citations: 0
✹ Influential: 0
📄 PDF
🀖 AI Summary
研究针对自回園囟像生成䞭KV猓存访问瓶颈问题通过分析泚意皀疏性特埁提出对角线感知皀疏泚意机制星著提高倄理速床和效率。
📝 Abstract
Autoregressive image generation has emerged as a paradigm for multimodal AI systems due to its compatibility with transformer-based LLM serving infrastructures. However, generating thousands of visual tokens per request makes decoding increasingly bottlenecked by KV cache accesses during attention computation. Sparse attention is particularly attractive for this workload because many visual generation applications tolerate moderate quality degradation in exchange for improved performance and efficiency. While sparse attention has been extensively explored for text-based LLM inference, it remains unclear whether its sparsity assumptions generalize effectively to autoregressive image generation. We present the first systematic characterization of attention sparsity in autoregressive image generation across diverse workloads and representative open-source models. Our analysis reveals several distinguishing properties, including a pronounced prefill-decode asymmetry, strong attention concentration on prompt and local tokens, and a unique diagonal attention sparsity pattern arising from the spatial locality of visual tokens. Motivated by these observations, we propose a diagonal-aware sparse attention mechanism that selectively skips KV entries along the diagonal attention direction within a recent window. Implemented on top of a GPU-based serving system using FlexGen, FlashAttention-2, and custom kernels, our approach achieves up to 3.1x throughput and 1.19x latency improvements with less than 2% quality degradation compared to dense inference.
Problem

Research questions and friction points this paper is trying to address.

Autoregressive Image Generation
Attention Sparsity
KV Cache Accesses
Innovation

Methods, ideas, or system contributions that make the work stand out.

diagonal-aware sparse attention
autoregressive image generation
KV cache access optimization
attention sparsity
D
Daeun Kim
KAIST
J
Junwha Hong
Agency for Defense Development
C
Changhun Oh
KAIST
Y
Yoonsung Kim
KAIST
Y
Yoonhyeong Lee
Seoul National University
Jongse Park
Jongse Park
Associate Professor; School of Computing; KAIST
Computer ArchitectureHW/SW CodesignAI SystemsAutonomous Systems