Self-Indexing Attention for Compression-Compatible Sparse Long-Context LLM Inference

📅 2026-08-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fragmented retrieval strategies between prefill and decoding phases in sparse long-context inference, along with the non-reusability of indices and their incompatibility with KV cache compression. We propose a training-free unified indexing framework that constructs reusable token-level indices based on shared transform-domain sign-magnitude representations, introducing the first mechanism enabling a single representation to serve both grouped prefill selection and decoding retrieval. By integrating 1-bit indexing with bitwise operation acceleration, the method achieves efficient retrieval compatible with low-bit KV compression without requiring additional metadata. Experiments demonstrate that at 5% attention density, performance closely approximates dense attention while accelerating prefill and decoding by 6.1× and 10.3×, respectively, with verified broad compatibility across diverse compression methods.
📝 Abstract
Sparse long-context inference requires efficient token retrieval in both prefill and decode. Existing methods often use different retrieval strategies for the two stages, preventing one retrieval representation from being reused throughout inference. We propose Self-Indexing Attention, a training-free framework built on a shared transform-domain sign-magnitude representation. The key signs provide a reusable token-level index for grouped prefill selection and decode retrieval, while the same representation remains compatible with external KV-cache compression without separate indexer metadata. This 1-bit index enables efficient retrieval through bitwise operations widely supported by modern accelerators. At 5% attention density, Self-Indexing Attention remains close to dense attention on LongBench and RULER and achieves up to 6.1x prefill and 10.3x decode attention-operator speedups. Experiments with TurboQuant and DeepSeekV4-Flash further demonstrate compatibility with low-bit KV-cache compression and pretrained sparse-attention indexers.
Problem

Research questions and friction points this paper is trying to address.

sparse long-context inference
token retrieval
KV-cache compression
prefill and decode
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Indexing Attention
Sparse Long-Context Inference
KV-Cache Compression
Sign-Magnitude Representation
Training-Free Framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xu Yang
Hunan University
J
Jiapeng Zhang
Hunan University
Y
Yuxin Chen
Tencent
F
Feiqiang Sun
Tencent
C
Chengguang Xu
Tencent
Feng Jin
Feng Jin
Tencent
Zhuo Tang
Zhuo Tang
Central South University