PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the resource redundancy caused by redundant KV cache computation in multi-LoRA agent systems, where direct cache reuse compromises role specificity. We propose a training-free KV cache sharing framework that introduces two novel mechanisms, PreLRShared and ReBaseShared. By precomputing caches via low-rank decomposition and reconstructing a neutral base cache, our approach eliminates redundant prefill phases while preserving distinct behavioral characteristics for each agent role. The method achieves precise cache sharing and error correction without additional training. Coupled with single-stream and concurrent inference optimization strategies, it accelerates time-to-first-token latency by 3.1× and improves throughput by 2.3×, with only a marginal accuracy degradation of 1.1 points.
📝 Abstract
Multi-LoRA agent systems enable efficient role specialization by sharing a common backbone model. However, each agent repeatedly processes the growing shared trajectory and constructs its own KV cache, introducing substantial memory and computation redundancy in long-horizon tasks. Existing KV cache sharing methods reduce this repeated prefill, but they either require additional training or architectural constraints or retain substantial model computation. Moreover, direct cache reuse causes the current agent to rely on cache states generated by the previous agent's adapter, weakening the role-specific behavior encoded by its own LoRA. We present PReCache, a training-free KV cache sharing framework with two designs, namely PreLRShared and ReBaseShared, that share the base cache computed using the pretrained weights and precompute a compact agent-specific low-rank (LR) cache. To remove repeated prefill, PreLRShared precomputes each agent's LR cache when the shared context is first processed, allowing the current agent to use its own LR cache without reprocessing context processed by previous agents. To improve sharing accuracy, ReBaseShared reconstructs the shared base cache from adapter-free hidden states, reducing the remaining error caused by the previous agent's adapted representation. To minimize its reconstruction cost, we propose two inference schemes tailored to single-stream inference and concurrent serving, performing the same reconstruction after each agent's turn or alongside its execution, respectively. Across multiple models and agent benchmarks, PreLRShared achieves up to a 3.1x TTFT speedup and a 2.3x improvement in per-request throughput over inference without KV cache sharing. ReBaseShared best preserves accuracy overall among the evaluated cache-sharing methods, with an average drop of only 1.1 points relative to inference without cache sharing.
Problem

Research questions and friction points this paper is trying to address.

Multi-LoRA agents
KV cache sharing
memory redundancy
computation redundancy
role-specific behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV Cache Sharing
Multi-LoRA Agents
Low-Rank Precomputation
Training-Free
Neutral Reconstruction