Preserving Provenance in Shared KV Caches for LLM Serving

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the cross-adapter conflicts and privacy leakage in shared KV caches caused by ignoring the computational context provenance. It presents the first systematic investigation of this "provenance-blind reuse" phenomenon and proposes a universal descriptor mechanism based on KV provenance contracts. By binding request and worker node metadata to cache keys, this approach achieves provenance-dimensional extension without requiring connector-specific logic. Furthermore, differential inspection and multi-path cache key binding are implemented atop vLLM and SGLang. The proposed method completely eliminates unsafe reuse, effectively defends against prompt inference attacks, restores model accuracy to 0.94, and incurs less than 0.34ms of additional latency.
📝 Abstract
Production LLM serving stacks combine an inference engine's local prefix cache with a shared KV-cache tier for fleet-wide reuse. The local cache distinguishes requests by adapter, weight configuration and sharing domain, but the shared tier may key entries only by token content and coarse model metadata. This boundary erases provenance and lets identical tokens under incompatible computational or sharing contexts collide. We call this composition gap provenance-blind reuse and present its first systematic study. A source audit of three vLLM connectors confirms the structural omission, while runtime experiments reproduce it across vLLM and two SGLang releases, 12 models from 7 families (0.5 B-32 B), and over 160 configurations. Cross-adapter collisions reduce accuracy from 0.94 to 0.64, incompatible KV representations reduce reasoning accuracy to zero, and salt omission enables 93% prompt identification from timing. We formalize the missing guarantee as the KV provenance contract: for a declared dimension registry, shared keys must be injective over computational and sharing provenance, with identities stable across workers. Any dimension with a stable identity can therefore be added without connector-specific key logic. A canonical descriptor binds per-request and per-worker provenance into lookup and store keys, while a differential checker detects dimensions that change KV state without changing the key. Implemented in vLLM and SGLang 0.5.20 across three cache paths, provenance binding eliminates unsafe reuse while preserving legitimate sharing. Hit-path latency changes remain within 0.34 ms and below run-to-run variation; retention grows with provenance diversity.
Problem

Research questions and friction points this paper is trying to address.

KV cache
provenance-blind reuse
LLM serving
cross-adapter collision
shared cache
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV Cache Provenance
Provenance Contract
Canonical Descriptor
Differential Checker
Shared KV Cache
🔎 Similar Papers
No similar papers found.