🤖 AI Summary
This study addresses the efficiency bottlenecks in long-context Transformer decoding arising from the quadratic complexity of self-attention and the linear growth of the KV cache. To this end, we propose a community-detection-based sparse attention method. Specifically, our approach constructs a token graph using prefill QK scores and partitions it into semantic communities to facilitate efficient retrieval. Furthermore, a local update rule assigns incoming tokens in constant time, eliminating global repartitioning and enabling training-free streaming decoding. Extensive experiments on Qwen3 and Llama-3.1 demonstrate that the proposed method achieves up to a 1.71× improvement in end-to-end generation throughput without compromising accuracy.
📝 Abstract
Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely on semantically coarse heuristics or expensive clustering that is difficult to update efficiently during decoding. We introduce CommunityKV, a framework that formulates sparse attention as a community detection problem. CommunityKV constructs a token graph from the $QK^T$ scores already computed during standard prefill, and partitions the graph into communities to enable retrieval of semantically coherent token groups. A local update rule assigns newly generated tokens to communities in constant time, enabling sparse retrieval throughout streaming decoding without global re-partitioning. We evaluate CommunityKV on Qwen3 and Llama-3.1 models across three long-context benchmarks. With one graph per query head, CommunityKV delivers up to $1.25\times$ the end-to-end generation throughput of dense attention, while query-group graph aggregation yields up to $1.71\times$ with comparable accuracy.