Efficient Agentic LLM Serving over SSD-based Sparse KV Storage

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of SSD read blocking on the critical path and inefficient fragmented access during sparse attention-based LLM inference. To this end, we propose Janus, a framework that leverages the model's intrinsic modules to predict KV cache demands for prefetching without requiring additional training. Furthermore, Janus optimizes storage access through computation-I/O overlap, read-write interference suppression, read merging, and sequential write packing techniques. Experimental results demonstrate that Janus reduces time-to-first-token latency by 1.57–3.69× and improves average throughput by 1.22–1.85×, while accelerating agent serving and preserving decoding efficiency.
📝 Abstract
Agentic sessions driven by Large language models (LLMs) often alternate between model inference and tool use, accumulating long histories across successive rounds. Serving these sessions efficiently requires reducing attention computation and retaining history key-value (KV) caches to avoid recomputation. Recently, frontier open-source LLMs adopt sparse attention to reduce computation by selecting only part of the history, while SSDs provide a cheaper alternative to CPU DRAM for storing KV caches. However, sparse KV selection depends on the ad hoc intermediate values during model inference, so it forces SSD reads to lie on the inference critical path. These reads are further slowed by fragmented accesses and read-write interference in SSDs. To address these challenges, we present Janus, an agentic serving framework for sparse attention LLMs with SSD-centric KV storage. Janus focuses on append prefill, which processes each round's newly added inputs and accounts for most history KV loading. To move SSD reads out of the critical path, Janus runs the model's own KV selection module on earlier intermediate values, predicting KV demand without additional training. The predicted reads overlap with model computation, and any prediction misses are fetched before attention executes to preserve model outputs. To improve SSD efficiency, Janus coalesces adjacent reads, packs scattered KV pages into sequential writes on the CPU, and limits background writes while reads are active. Across three models and three agentic traces, Janus outperforms existing works by up to 1.57-3.69 times (1.22-1.85 times on average) in terms of the time to first token latency, while maintaining decode efficiency.
Problem

Research questions and friction points this paper is trying to address.

Agentic LLM Serving
Sparse Attention
KV Cache
SSD Storage
Critical Path
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Attention
SSD-based KV Storage
Agentic LLM Serving
Append Prefill
Critical Path Optimization
🔎 Similar Papers
No similar papers found.