Characterizing High Bandwidth Flash for LLM Serving
This study addresses the memory capacity and bandwidth bottlenecks in large language model (LLM) serving, along with the challenges of KV cache reuse under agentic workloads. To this end, it proposes an HBM-HBF host-side tiered storage architecture leveraging high-bandwidth flash (HBF). By designing a buffer-cache-aware scheduling strategy to coordinate data placement and incorporating real-trace simulation techniques, the system achieves holistic optimization across performance, energy consumption, and device endurance. Experimental results demonstrate that the proposed architecture reduces task completion time by 36.1%–87.0%, yields 55.8% energy savings, and significantly extends the HBF write lifetime from 4.77 to 14.82 years.