Score
Design, build, and evaluate methods for mapping and transforming key-value (kv) or dense latent cache representations from a sender model to a receiver model so the receiver preserves contextual knowledge, observations, and intermediate reasoning signals. This includes creating dense latent alignment and cache-transfer procedures that operate directly on cache tensors (latent states) rather than decoding to and re-encoding from text.
This work addresses the challenges of heterogeneous multi-agent communication—namely, substantial information loss, high computational overhead, and misalignment in latent spaces—by proposing a dense KV-cache communication method tailored for heterogeneous agents. The approach leverages lightweight cross-model cache transformation and a two-stage training scheme (reconstruction followed by generation) to jointly transfer both perceptual content and reasoning logic. It reveals, for the first time, a duality in information structure between context-aware and context-agnostic transmission, enabling efficient “mind-reading” communication without requiring shared inputs or homogeneity assumptions. Evaluated across six heterogeneous agent pairs based on the Qwen3 series and six in- and out-of-domain benchmarks, the method matches or surpasses text-based communication while using only half to one-third of the computational cost, and remains effective even when receivers lack contextual input, significantly outperforming existing heterogeneous baselines.
This work addresses the high latency and information loss inherent in text-based communication among large language model agents, as well as the limited adaptability of existing cache-transfer methods to contextual discrepancies. To overcome these challenges, the authors propose a text-free cross-model communication mechanism that leverages a lightweight adapter (only 13 MB) to directly transmit compressed key-value (KV) cache summaries, enabling efficient state sharing. The core innovations include a joint compression-and-translation framework for KV caches, a context-aware summarization strategy, and a highly compact adapter architecture. Experimental results demonstrate that the proposed method outperforms the 956 MB C2C model in identical-context scenarios and achieves a 23% improvement in communication accuracy alongside an 8.5× speedup in cross-context settings.
Large language models (LLMs) suffer from quadratic attention complexity and explosive KV cache memory consumption and latency when processing long contexts (≥32K tokens). To address this, we propose the first comprehensive methodology classification framework for KV cache compression, unifying key techniques—including quantization, pruning, low-rank approximation, sequence grouping, and locality-aware reuse—across both theoretical principles and practical implementation dimensions. Through standardized cross-method evaluation, we characterize the fundamental trade-offs among accuracy, inference speed, and memory footprint. Experimental results demonstrate that, with perplexity degradation under 1%, the optimal compression strategy achieves 2–5× KV cache memory reduction and 1.3–2.1× inference speedup, substantially enhancing the feasibility of deploying LLMs in long-context scenarios.
To address the high memory overhead and low retrieval accuracy of KV caches in long-context LLM inference, this paper proposes a retrievable KV cache compression method based on semantic clustering. Unlike conventional position-based paging or irreversible pruning strategies, our approach manages caches at the granularity of semantic clusters, integrating dynamic cache selection, lightweight indexing, and hierarchical cache management to enable fine-grained, high-fidelity real-time semantic retrieval and reconstruction. Evaluated on 32K-context workloads, the method achieves negligible accuracy degradation while operating under tight KV cache budgets of only 1K–2K entries. It reduces inference latency by 2× and improves decoding throughput by 2.5×, significantly outperforming existing retrievable compression approaches.
To address the GPU memory bottleneck induced by KV cache in long-context LLM inference, this work pioneers the application of Product Quantization (PQ) to KV cache compression, framing it as an approximate nearest neighbor search in embedding space. We propose an overlapping block partitioning scheme coupled with a hierarchical caching mechanism that eliminates extraneous computation and communication overhead across both prefill and decode stages. Our method preserves model quality while substantially reducing service latency: it achieves a 4.60% score improvement on InfiniteBench, outperforms state-of-the-art methods in both prefill and decode latency, and enables efficient inference over context lengths exceeding 10,000 tokens. The core contributions are (i) a PQ-driven KV compression paradigm that drastically reduces memory footprint without accuracy degradation, and (ii) a zero-overhead hierarchical scheduling design that seamlessly integrates compression into the inference pipeline.
This study addresses a critical limitation in existing latent communication frameworks, where receivers fail to substantively utilize message content and shared models are easily supplanted by irrelevant inputs. To overcome this, we propose a side-memory communication mechanism based on draft key-value (KV) states, integrating linear projection, gated attention, and a progressive training strategy to achieve efficient, content-sensitive communication with frozen base models. This work pioneers the draft KV communication paradigm, requiring only 1/348 of the parameters used by C2C while ensuring that transmitted messages remain indispensable. Empirical evaluations demonstrate that our approach attains 78.04% accuracy on MMLU-Redux, with performance scaling significantly alongside the size of the shared model. Furthermore, the proposed method exhibits robust cross-task transferability, highlighting its potential for scalable and parameter-efficient multi-agent collaboration.
为解决长上下文推理中KV缓存内存需求大的问题,PuzzleKV通过将每个KV缓存分块并进行低秩分解以压缩存储成本,同时保持高精度。
This study addresses the memory surge and high latency challenges in long-context reasoning for large language models, noting that existing KV cache compression methods overlook the critical impact of relative distance on retrieval capability. We propose Distance-KV, which reveals for the first time that relative distance constitutes a core structural dimension for long-context retrieval. This method learns a static KV retention pattern by jointly optimizing across layers, attention heads, and relative distances offline on a frozen language model, enabling direct cache pruning during inference without online importance scoring. Experimental results demonstrate that Distance-KV outperforms the strongest baseline by 9.3 points on the RULER benchmark while reducing the memory footprint of Llama-3.1-8B by 65.4% and accelerating decoding speed by 1.66×.
This study addresses the bottleneck of excessive memory and context overhead in KV cache transmission within multi-agent systems, which frequently exceeds GPU resource limits. To this end, it proposes a "receiver-conditioned communication" paradigm and introduces CacheBack, a training-free method that filters and compresses sender-side KV caches based on attention weight analysis. By transmitting only the information required by the receiver, CacheBack enables precise, on-demand communication across heterogeneous architectures such as Transformers and Mamba. Experimental results demonstrate that the proposed approach eliminates 75% of redundant states, improves accuracy by 14.7%, and reduces latency by 3.2×, while exhibiting consistent generalization across multiple model families.
本文提出MetaKV框架,通过自适应选择KV缓存压缩配置来解决大语言模型推理中的内存开销问题,同时满足用户指定的延迟和峰值内存预算。