🤖 AI Summary
This study addresses the KV cache memory bandwidth bottleneck in long-context decoding, where existing compression methods suffer from error accumulation due to decoupled retention and compensation mechanisms. To overcome this, it proposes a unified framework grounded in exact factorization, leveraging calibration distributions to simultaneously drive complementary KV state retention and evicted quality reallocation. A novel coverage calibration mechanism is introduced to unify retention and writing channels, eliminating reliance on independent predictors or online determinant computations. Furthermore, efficient management is achieved through offline allocation distillation, lightweight cache-aware indexing, and recursive query-adaptive normalization. Empirical evaluations demonstrate that the proposed method improves performance on the RULER benchmark by 3.78 points at a 90% compression rate while achieving strong results on LongBench, significantly enhancing both decoding efficiency and effectiveness.
📝 Abstract
Long-context decoding is increasingly constrained by key--value (KV) cache memory and bandwidth. Existing fixed-budget compression methods typically separate retention from compensation, while a retention ranking specifies neither discarded attention mass nor the direction of induced output error. We start from an exact factorization: eviction error equals evicted attention mass times the directional gap between the evicted centroid and retained output, highlighting the importance of set-level coverage in retention and mass-preserving memory writing. We introduce CORE COverage Calibration and Evicted-Mass REdistribution for KV Cache, which distills an offline allocation combining query utility and log-determinant coverage into a lightweight cache-aware indexer. At inference, one calibrated distribution drives both channels: its Top-$B$ ordering retains complementary KV states, while its excluded allocation mass and conditional weights parameterize latent-memory writes without a separate write-weight predictor or online log-determinant evaluation. Our analysis provides a four-term pre-compensation error certificate, characterizes non-additive coverage interactions, and establishes mass-independent write stability with hierarchical bounds through recurrent and query-adaptive normalization. Across three backbones, CORE exceeds the strongest RULER baseline by up to 3.78 points at 90\% compression; LongBench and repeated-eviction evaluations further demonstrate strong effectiveness and decoding efficiency.