SPIN: Shadow Predictive Indexer for Sparse Attention

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational bottleneck in long-context inference caused by exhaustive scoring in sparse attention indexers. To overcome this limitation, we propose SPIN, a framework that integrates KV block management with speculative decoding. By leveraging lightweight historical predictions to substitute layer-wise full-cache scoring, SPIN precisely identifies critical KV blocks while avoiding redundant computation. The proposed framework has been integrated into the vLLM serving engine. Experimental results demonstrate that SPIN achieves 30%–40% attention sparsity without compromising task performance, yielding a 14.9% improvement in throughput and a 13.2% reduction in latency. This work establishes a new paradigm for efficient long-sequence inference.
📝 Abstract
Indexer-based sparse attention reduces the cost of core attention by passing only a fixed, small number of important tokens to it. However, the indexer must still score the entire KV cache at every decoding step. This scoring overhead becomes a major bottleneck as the context length grows. We propose SPIN (Shadow Predictive Indexer) to reduce this indexer overhead. SPIN uses lightweight, history-based prediction to identify important KV blocks, avoiding the need to score the full KV cache at every decoding step. SPIN treats KV blocks and speculative decoding as first-class design and implementation considerations. Across extensive evaluations on long-context and agentic benchmarks, SPIN achieves 30-40% sparsity while preserving task quality. In end-to-end vLLM serving, SPIN improves output throughput by up to 14.9% and reduces median inter-token latency by up to 13.2%.
Problem

Research questions and friction points this paper is trying to address.

Sparse Attention
Indexer Overhead
KV Cache
Long Context
Decoding Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Attention
Predictive Indexer
KV Cache
Speculative Decoding
Long-context
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.