🤖 AI Summary
This study addresses the prohibitive memory overhead of key-value (KV) caching during long-context inference in large language models, noting that existing compression methods typically require architectural modifications and incur substantial computational costs. To overcome these limitations, this work proposes a symmetry-aware value cache merging strategy. By synergistically integrating cross-layer symmetry analysis, a value cache merging algorithm, and orthogonal compression techniques, the approach achieves efficient KV cache reduction without altering the underlying model architecture. The primary contribution lies in significantly decreasing cache memory consumption with minimal additional overhead, thereby enabling context window extension in memory-constrained scenarios while preserving stable decoding performance.
📝 Abstract
While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths. One solution is to compress this memory by exploiting inter-layer cache similarities. However, most existing techniques necessitate architectural changes to LLMs and incur substantial overhead. In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding. Furthermore, we show that this approach can be exploited alongside existing cache compression techniques, composing with high-ratio quantization or key cache pruning to reach compression ratios that neither method reaches alone, with minimal additional cost. Ultimately, our findings reveal a major source of underutilized capacity in the value cache, offering a simple yet highly effective direction for scaling context windows under memory constraints.