PatchKV: Weight-Space Compensation of KV Cache

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the linear growth of KV cache memory during long-context inference in large language models, where aggressive compression often causes severe performance degradation. To overcome this, we propose a training-free weight patch compensation framework that pioneers encoding partial context information directly into model weights, transcending conventional cache-only optimization paradigms. By leveraging closed-form ridge regression solutions, block-level activation alignment, and weight-space patching techniques, the compensation is computed and merged during model loading. This approach enables high-fidelity long-text generation with zero additional inference overhead. Extensive evaluations across multiple benchmarks demonstrate substantial improvements in compression performance under aggressive budgets, establishing a new paradigm for efficient long-context inference.
📝 Abstract
Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade sharply at aggressive compression budgets. We propose PatchKV, a training-free framework that compensates KV cache compression methods by carrying part of the context in the model's weights. PatchKV pairs an off-the-shelf compressed KV cache with a context-specific weight patch, which is computed once at context-loading time and served for downstream queries for the context. The weight patch is derived in closed form via ridge regression, by aligning the block-wise activations of context-derived reference query tokens under the full cache and the compressed cache. Once merged into the model, the patch leaves the forward graph and per-query inference cost unchanged in the single-context, multi-query setting. Across long-context QA (SCBench with up to 170K tokens, SQuAD, NIAH) and math (GSM8K) benchmarks on three model architectures, PatchKV consistently improves cache compression methods, suggesting an alternative direction to compensate them at aggressive budgets.
Problem

Research questions and friction points this paper is trying to address.

KV cache compression
long-context inference
memory bottleneck
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV cache compression
training-free
weight patch
ridge regression
long-context inference
🔎 Similar Papers
No similar papers found.