HCRMap: Pressure-Aware Hot-Expert Residency Mapping for 3.5D MoE Chiplet Inference

📅 2026-07-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the imbalance in computation, communication, memory bandwidth, and execution queues caused by uneven expert popularity during Mixture-of-Experts (MoE) large model inference in 3.5D multi-chiplet systems. To tackle this challenge, the paper proposes a dynamic hot-expert residency mapping framework that introduces, for the first time in 3.5D MoE inference, a pressure-aware dynamic expert replication mechanism. By jointly modeling expert popularity, weight loading overhead, migration cost, and runtime resource pressure, the framework intelligently determines expert replica promotion/demotion, placement, and token routing. Through efficient cross-memory-hierarchy reuse of hot experts and improved load balancing, the proposed approach reduces end-to-end latency by up to 46.7% during the prefill phase and 46.0% during the decode phase compared to state-of-the-art methods such as Hydra, MoEntwine, and PIMoE.
📝 Abstract
Mixture-of-Experts (MoE) large language models (LLM) activate only a small number of experts during inference, but token routing introduces persistent expert hotness skew: a small set of hot experts continuously receives most tokens, while the remaining experts are lightly loaded. On 3.5D multi-chiplet systems, this skew not only causes compute imbalance but also amplifies pressure on communication, memory bandwidth, I/O, and execution queues. Therefore, the core problem is not simply to reduce token movement, but to dynamically place and reuse hot expert replicas across different memory tiers. This paper proposes HCRMap, a hot expert residency mapping framework for pressure-aware expert replica management in 3.5D MoE inference. Based on expert hotness, weight loading cost, migration overhead, and runtime resource pressure, HCRMap dynamically determines which experts should be promoted, retained, demoted, or evicted. It then maps routed token groups to suitable resident replicas, thereby jointly mitigating communication, memory, and queue bottlenecks. Experimental results show that HCRMap reduces end-to-end latency by 43.6% and 43.0% over Hydra in the prefill and decode stages, respectively; by 34.5% and 33.1% over MoEntwine; and by 46.7% and 46.0% over PIMoE.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
expert hotness skew
3.5D chiplet
memory hierarchy
resource pressure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hot-expert residency
Pressure-aware mapping
3.5D chiplet
Mixture-of-Experts
Dynamic replica management
🔎 Similar Papers