Recall Before You Rank: Similarity-Guided Top-$K$ Reuse for Efficient Long-Context Attention

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the computational inefficiency of dynamic Top-K sparse attention in long-context decoding, where scoring the full key-value (KV) cache and performing global Top-K selection incurs linearly growing overhead with context length. To mitigate this, the authors propose ReTopK, a training-free acceleration method that retrieves historical queries most similar to the current query, reuses their support sets, and merges them with a recent local window to form a compact candidate set on which scores are re-ranked. ReTopK incorporates similarity-based support set reuse, local re-ranking, similarity-triggered fallback to full history, and periodic cache refreshing, while storing only indices to minimize memory overhead. Experiments show that across context lengths from 16K to 128K, ReTopK outperforms existing approximation methods on PG19 perplexity, NIAH, and LongBench; at 128K tokens, it achieves only a 0.50% perplexity increase while accelerating attention computation by 3.07×.
📝 Abstract
Top-$K$ sparse attention reduces the cost of Softmax and value aggregation by attending to only a small subset of key--value (KV) entries. However, identifying this subset still requires scoring the current query against the full KV cache and performing global Top-$K$ selection, leaving selector cost linear in context length and limiting the practical efficiency of sparse attention for long-context decoding. In this paper, we introduce ReTopK, a training-free method that accelerates dynamic Top-$K$ attention by reusing historical retrieval decisions. ReTopK builds on the observation that similar queries often attend to overlapping supports and that partially overlapping supports can still preserve most of the Exact Top-$K$ attention mass. For each attention head, it maintains a bounded cache of historical query--support pairs, retrieves the most similar cached queries for each new query, unions their stored supports with a recent window, and reranks only the resulting compact candidate set using exact current-query scores. A similarity-based fallback invokes full-history Exact Top-$K$ when reuse is unreliable, while periodic exact refreshes limit cache drift. ReTopK retains the complete KV cache and reuses only selected indices, rather than historical scores, attention weights, or outputs. Across 16K--128K contexts, ReTopK achieves the lowest PG19 perplexity and the highest NIAH and LongBench scores among the evaluated approximate methods. At 128K with $K=512$, ReTopK incurs only a 0.50\% perplexity increase over Exact Top-$K$ while accelerating attention computation by $3.07\times$.
Problem

Research questions and friction points this paper is trying to address.

long-context attention
Top-K sparse attention
KV cache
attention efficiency
context length
Innovation

Methods, ideas, or system contributions that make the work stand out.

ReTopK
Top-K sparse attention
long-context attention
similarity-guided reuse
training-free acceleration
W
Wenshuai Yao
School of Integrated Circuits, Peking University, Beijing, China
Wenyong Zhou
Wenyong Zhou
The University of Hong Kong
Computer Vision
H
Hanyong Shao
School of Integrated Circuits, Peking University, Beijing, China
Y
Yizhe Chen
School of Integrated Circuit Science and Engineering, Beihang University, Beijing, China
Zhiyuan Ning
Zhiyuan Ning
Westlake University
Graph Machine LearningKnowledge GraphsLarge Language Models
Y
Yuannuo Feng
School of Integrated Circuit Science and Engineering, Beihang University, Beijing, China
R
Ru Huang
School of Integrated Circuits, Peking University, Beijing, China
Kechao Tang
Kechao Tang
Peking University
HfO2 ferroelectricsmemory devicesvanadium oxide