CommunityKV: Efficient Long-Context Decoding via Graph Partitioning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the efficiency bottlenecks in long-context Transformer decoding arising from the quadratic complexity of self-attention and the linear growth of the KV cache. To this end, we propose a community-detection-based sparse attention method. Specifically, our approach constructs a token graph using prefill QK scores and partitions it into semantic communities to facilitate efficient retrieval. Furthermore, a local update rule assigns incoming tokens in constant time, eliminating global repartitioning and enabling training-free streaming decoding. Extensive experiments on Qwen3 and Llama-3.1 demonstrate that the proposed method achieves up to a 1.71× improvement in end-to-end generation throughput without compromising accuracy.
📝 Abstract
Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely on semantically coarse heuristics or expensive clustering that is difficult to update efficiently during decoding. We introduce CommunityKV, a framework that formulates sparse attention as a community detection problem. CommunityKV constructs a token graph from the $QK^T$ scores already computed during standard prefill, and partitions the graph into communities to enable retrieval of semantically coherent token groups. A local update rule assigns newly generated tokens to communities in constant time, enabling sparse retrieval throughout streaming decoding without global re-partitioning. We evaluate CommunityKV on Qwen3 and Llama-3.1 models across three long-context benchmarks. With one graph per query head, CommunityKV delivers up to $1.25\times$ the end-to-end generation throughput of dense attention, while query-group graph aggregation yields up to $1.71\times$ with comparable accuracy.
Problem

Research questions and friction points this paper is trying to address.

Long-context decoding
Sparse attention
KV cache
Self-attention
Transformers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Attention
Community Detection
Graph Partitioning
KV Cache
Long-Context Decoding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.