ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the linear growth of KV caches and the computational overhead of reconstruction anchor search during long-context LLM inference by proposing a selective amortization principle. Specifically, it trains a value-aware indexer to select anchors in a single pass and amortize them across contexts. By integrating convex-hull-constrained key merging with attention bias fitting, the method enables a cache to be constructed once and reused multiple times, preserving context-specific reconstruction while significantly reducing overhead. To our knowledge, this is the first approach to amortize anchor selection strategies across contexts. Empirical evaluations on benchmarks such as QuALITY demonstrate superiority over existing methods; at a 10% retention rate, the proposed approach achieves an accuracy of 0.6474 while reducing compression time approximately 25-fold to 37.3 seconds.
📝 Abstract
Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods such as Attention Matching achieve strong downstream task performance with compact KV caches. However, iterative anchor search dominates the compaction cost of OMP-based Attention Matching. This motivates our selective amortization principle of learning a reusable anchor-selection policy across contexts while retaining context-specific reconstruction. In this work, we propose ARC-KV, a novel reconstruction-based KV cache compaction method that follows this principle. To this end, we first train a value-aware indexer to select real-key anchors in a single scoring pass. ARC-KV then applies convex-hull-constrained key merging and fits an attention-mass bias and compact values against the full cache. At inference time, ARC-KV builds the compact cache once per context using the frozen indexer and reuses it for all subsequent queries. Extensive experiments demonstrate that ARC-KV outperforms reported compaction methods in most settings across QuALITY, RULER, and LongBench on Llama-3.1-8B-Instruct. In particular, at 10% KV retention on QuALITY, ARC-KV improves accuracy from 0.6409 to 0.6474 over Attention Matching while reducing compaction time by a factor of 25.73, from 959.8 s to 37.3 s.
Problem

Research questions and friction points this paper is trying to address.

KV cache compaction
long-context inference
anchor search overhead
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV Cache Compaction
Reconstruction-Based Attention Matching
Value-Aware Indexer
Selective Amortization
Long-Context LLM Inference
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.