More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory and computational bottlenecks inherent in attention mechanisms during autoregressive decoding with large language models, where query-key matching often becomes a critical constraint following sparsification. To overcome this, the authors propose the SAGA architecture, which decouples the number of key and value heads, integrates Atop-N approximate sparse attention, and introduces an efficient fine-tuning strategy enabling seamless adaptation of pretrained models. The work theoretically establishes the advantages of asymmetric head configurations. Empirically, the proposed approach achieves over twofold decoding acceleration in long-context scenarios, while training-from-scratch performance closely approximates Grouped Query Attention baselines. Furthermore, it facilitates low-cost conversion of existing models, offering a practical and theoretically grounded solution for efficient long-context inference.
📝 Abstract
utoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitigate this cost by selecting only high-probability entries of the attention matrix. We observe that in many such methods, this renders the probability-value multiplication negligible, shifting the bottleneck to the query-key step. Key heads can therefore be reduced to accelerate inference, while retaining more value heads preserves capacity with limited additional decoding cost. We introduce Sparse Asymmetric Group-Query Attention (SAGA), which decouples key and value head counts to exploit this principle, and pair it with approximate top-N (Atop-N) attention, a simple sparse attention method designed to study the interaction between sparsity and head-count asymmetry. We formalize the benefits of this asymmetry theoretically and validate them empirically through latency measurements and quality evaluations on models up to 1.5B parameters. Together, SAGA and Atop-N achieve end-to-end decoding speedups exceeding $2\times$ over our full-attention GQA baseline at long contexts. Models trained from scratch with SAGA nearly match the quality of comparable GQA variants on the evaluated benchmarks. To facilitate adoption, we introduce an efficient fine-tuning method that converts pretrained models to the SAGA architecture, enabling practitioners to benefit from our approach without costly retraining.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Sparse Attention
Autoregressive Decoding
Inference Acceleration
Grouped-Query Attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Asymmetric Attention
Group-Query Attention
Approximate Top-N Attention
LLM Decoding Acceleration
Efficient Fine-tuning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Noam Elata
Technion – Haifa, Israel
I
Itay Lamprecht
Technion – Haifa, Israel; Crusoe AI
M
Mikey Shechter
Technion – Haifa, Israel
Daniel Ohayon
Daniel Ohayon
Technion University
I
Itay Hubara
Stealth Startup
Daniel Soudry
Daniel Soudry
Associate Professor
Neural NetworksMachine LearningTheoretical neuroscience