Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory surge and high latency challenges in long-context reasoning for large language models, noting that existing KV cache compression methods overlook the critical impact of relative distance on retrieval capability. We propose Distance-KV, which reveals for the first time that relative distance constitutes a core structural dimension for long-context retrieval. This method learns a static KV retention pattern by jointly optimizing across layers, attention heads, and relative distances offline on a frozen language model, enabling direct cache pruning during inference without online importance scoring. Experimental results demonstrate that Distance-KV outperforms the strongest baseline by 9.3 points on the RULER benchmark while reducing the memory footprint of Llama-3.1-8B by 65.4% and accelerating decoding speed by 1.66×.
📝 Abstract
The memory usage and decoding latency of LLM inference grow rapidly with context length. To reduce these costs, key-value (KV) cache compression methods selectively retain cached states based on token importance or differences in attention patterns across heads. However, we discover that retrieval capability varies substantially with relative distance, even within the same attention head. To exploit this structure, we introduce Distance-KV, which learns a static KV retention pattern over the joint space of layers, attention heads, and relative distances. The pattern is learned offline with the language model frozen and reused across inputs to prune and compact the KV cache without online importance scoring. Across three backbone models and four long-context benchmarks, Distance-KV consistently achieves the best overall performance among competing KV cache compression methods, exceeding the strongest compression baseline by up to 9.3 points on RULER at 128K. On Llama-3.1-8B-Instruct at 128K, Distance-KV reduces KV cache memory by 65.4% and achieves a $1.66\times$ decoding speedup relative to Dense. Together, these results identify relative distance as an important structural dimension for understanding how LLMs retrieve information over long contexts and for designing more efficient inference methods.
Problem

Research questions and friction points this paper is trying to address.

Long-context inference
KV cache compression
Relative distance
Memory efficiency
Decoding latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV cache compression
relative distance
long-context inference
static retention pattern
offline learning