🤖 AI Summary
This study addresses the GPU memory bottlenecks and inference inefficiencies caused by the linear growth of KV caches in many-shot in-context learning. To this end, it proposes MILO, a framework that introduces a pioneering block-granularity low-rank compression strategy exploiting the inherent low-rank redundancy within KV caches. Furthermore, MILO incorporates an information entropy-driven dynamic rank allocation algorithm that adaptively adjusts rank budgets across heterogeneous contexts to balance accuracy and efficiency. Experiments conducted on Qwen2.5 demonstrate that MILO reduces KV cache memory consumption by 50% and improves throughput by 1.8×, while maintaining nearly lossless performance on both classification and reasoning tasks.
📝 Abstract
Many-shot in-context learning (ICL) enables large language models (LLMs) to adapt to complex tasks by conditioning on thousands of demonstration examples, but this paradigm shifts the inference efficiency bottleneck to the key-value (KV) cache memory. Due to the linear scaling behavior of the KV cache, storing these intermediate tensors has become a paramount challenge for both online serving and on-device deployment. To address this issue, we propose a novel compression framework, termed MILO, that exploits the low-rank redundancy inherent in many-shot contexts. Specifically, MILO features a block-wise low-rank compression strategy that compresses the KV cache at the block granularity, where each block contains multiple many-shot examples. Furthermore, to handle the heterogeneous context density across different blocks, MILO dynamically allocates rank budgets based on the information entropy, preserving the fidelity of critical blocks while aggressively compressing redundant ones. Experimental results on Qwen2.5 models demonstrate that our method achieves up to 50% reduction in KV cache memory and 1.8x throughput improvement, with negligible performance degradation on classification and reasoning benchmarks, significantly outperforming prior baselines.