🤖 AI Summary
This study addresses the unclear efficacy of prefix cache replacement policies under LLM agent workloads, where complex algorithms frequently underperform. Through production trace analysis, it reveals the structural reasons why LRU proves optimal due to session regularity. The work proposes a compute-saving ratio metric and designs a hybrid lightweight policy combining rapid demotion with compute-aware partial eviction, further optimizing HBM-constrained scenarios via capacity-dependent granularity control. This research validates the effectiveness of recency-based strategies and elucidates the fundamental reasons why sophisticated approaches offer no benefit. Additionally, it open-sources the associated trace datasets and simulator, providing foundational support for future research in this domain.
📝 Abstract
Long-running LLM applications repeatedly send growing context, making prefix caching critical for reducing prefill cost. Yet prefix-cache behavior under agentic workloads remains poorly understood. We study production traces from two companies and evaluate 14 eviction algorithms across HBM-constrained and large memory-pool settings. Despite a large gap to Belady, sophisticated policies designed for traditional caches provide little benefit over LRU. The reason is structural: prefix reuse is dominated by the regular pacing of active sessions, making recency unusually predictive. Prefix caching nevertheless introduces new challenges, including heavy-tailed session footprints and highly variable miss costs as attention computation grows with sequence length. We introduce the compute-savings ratio and two offline oracles to quantify these effects. Our results show that effective prefix-cache management should retain recency as its foundation while selectively adding quick demotion for one-hit prefixes, compute-aware partial eviction for expensive misses, and capacity-dependent eviction granularity. We will release the traces and simulator to support future research.