Beyond Capacity: Scalable MoE LLM Inference via High-Bandwidth Flash with Direct GPU and HBM Paths

📅 2026-08-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory bottlenecks and low bandwidth utilization of existing flash architectures that constrain Mixture-of-Experts (MoE) inference. We propose a dual-path concurrent expert weight transfer architecture featuring a novel direct-and-relay path coordination mechanism. By integrating proactive expert prediction with KV cache traffic isolation, this approach effectively eliminates shared bottlenecks and masks storage latency. Experimental evaluations on typical MoE workloads demonstrate that the proposed system achieves 1.94× higher throughput and 1.90× end-to-end speedup compared to conventional designs. These results confirm the efficacy of concurrent heterogeneous storage access in significantly overcoming performance limitations for large-scale MoE inference.
📝 Abstract
Modern mixture-of-experts (MoE) language models increasingly strain the capacity and cost efficiency of high-bandwidth memory (HBM), as rapidly growing expert weights must be provisioned close to GPUs. High-bandwidth flash (HBF) offers substantially greater capacity, but conventional designs typically deliver HBF-resident expert weights to the GPU through HBM, leaving an additional direct GPU-HBF connection underutilized. We explore an HBF organization that simultaneously exploits two independent expert-delivery routes: a direct path that transfers expert weights from HBF to the GPU and a relay path that transfers them from HBF through the HBM base die to the GPU. Whole experts are assigned to one of the two routes, and transfers over both routes proceed concurrently, increasing aggregate expert-delivery bandwidth without replicating expert weights or introducing a shared relay bottleneck. Early expert determination identifies upcoming experts ahead of their conventional execution point, allowing HBF read latency to overlap with preceding computation, while separate management of immutable expert weights and mutable KV-cache data reduces interference between the two traffic classes. We evaluate the architecture using an event-driven continuous-batching LLM serving simulator with empirically measured GPU compute latencies. Across representative MoE workloads, concurrently utilizing the direct GPU-HBF and HBF-HBM-GPU routes consistently improves expert-delivery efficiency over designs restricted to either route alone. For a representative workload, the proposed architecture can achieve 1.94$\times$ higher throughput and 1.90$\times$ end-to-end speedup over a design that delivers all HBF-resident expert weights to the GPU through the HBM base die.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
LLM Inference
High-Bandwidth Flash
Memory Capacity
Expert Delivery Bandwidth
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
High-Bandwidth Flash
Dual-Path Data Delivery
Early Expert Determination
LLM Inference
🔎 Similar Papers
No similar papers found.