🤖 AI Summary
This study addresses the linear growth of KV cache during decoding in long chain-of-thought reasoning, where existing eviction strategies risk discarding critical historical information. We propose an online low-rank compression scheme that preserves a recent context window while folding historical states via block-incremental singular value decomposition and dynamic low-rank representations. Furthermore, a coordinate synchronization mechanism is introduced to eliminate historical drift caused by basis updates, achieving position-consistent compression throughout the sequence. Evaluated on DeepSeek and Qwen models, our method attains 4–5× KV cache compression with accuracy closely matching the original models and significantly outperforming existing baselines.
📝 Abstract
Long chain-of-thought reasoning substantially increases KV-cache memory during autoregressive decoding, as every generated token introduces new key and value states and causes the cache to grow linearly with decoding length. Existing KV-cache compression methods typically control this growth through token eviction, but irreversible deletion can remove historical states that later reasoning may need to revisit. SVD-based low-rank compression provides an alternative by retaining all positions with a more compact representation. However, extending it from a fixed prompt cache to online decoding is non-trivial. Through our investigation, we find that if the basis is updated for new tokens while old tokens keep their coordinates in the old basis, the stored history drifts substantially. Based on this observation, we propose iS-KV, an online low-rank KV-cache compression method for long-horizon reasoning. iS-KV keeps a recent window exact while incrementally folding older states into bounded-rank representations. As the low-rank basis evolves, it synchronizes historical coordinates with the updated basis to maintain representation consistency. On DeepSeek-R1-Distill-Llama-8B, iS-KV achieves 82.6% accuracy at 4.06-fold persistent-KV compression, close to the original model's 83.6%. On Qwen3-8B, it achieves 89.2% accuracy at 5.64-fold compression. Under matched memory budgets, iS-KV consistently outperforms token-eviction baselines.