MoSE: Mode-Switching Expander for Mixed LLM Training and Inference

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of static network topologies in accommodating both inference key-value (KV) transmission and training global connectivity under mixed workloads. To overcome this, the authors propose a reconfigurable expander that formulates topology design as a fixed-degree edge allocation problem. A novel coarse-grained mode-switching mechanism is introduced to dynamically adapt between dual modes without requiring additional ports or routing modifications. This approach integrates flow-level topology modeling, shortest-path routing, and random regular expander graph techniques to achieve efficient connectivity. Experimental results demonstrate that the proposed method reduces KV communication costs by over 90% in inference mode and decreases training communication overhead by approximately 27%, highlighting its effectiveness for heterogeneous workloads.
📝 Abstract
AI clusters increasingly run large language model (LLM) inference and training on the same fabric. Prefill-decode (P-D) disaggregation creates key-value (KV) cache transfers between prefill and decode groups, whereas training collectives and all-to-all traffic benefit from near-uniform global connectivity. A static sparse topology can therefore be poorly matched to one of the two traffic patterns. We present Mode-Switching Expander (MoSE), a reconfigurable expander that treats topology design as a fixed-degree edge-allocation problem. MoSE reallocates the same sparse edge budget toward direct P-D connectivity in inference-heavy modes and restores a uniform random regular expander in training-heavy modes. We evaluate MoSE using a 1024-group flow-level topology model, shortest-path routing, and two mixed workloads. Across 20 seeds, MoSE reduces average and 95th-percentile (P95) load-aware KV communication cost by 90.8\% and 91.9\% relative to Static-Training in the inference-heavy mode. In the training-heavy mode, it reduces average and P95 training communication cost by 22.7\% and 27.6\% relative to stale Static-Inference. These results show that coarse-grained topology switching can support both workload modes without additional ports or routing changes.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Mixed Workloads
Network Topology
Prefill-Decode Disaggregation
AI Clusters
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mode-Switching Expander
Reconfigurable Topology
Prefill-Decode Disaggregation
Mixed Workloads
Edge-Allocation
🔎 Similar Papers
No similar papers found.
F
Fan Yang
Southern University of Science and Technology; Faculty of Computer Science and Artificial Intelligence, Shenzhen University of Advanced Technology
Y
Ying Zhou
School of Electronic and Information Engineering, Beijing Jiaotong University
B
Binglei Wang
Southern University of Science and Technology; Faculty of Computer Science and Artificial Intelligence, Shenzhen University of Advanced Technology
Z
Zhenjie Zhou
Southern University of Science and Technology; Faculty of Computer Science and Artificial Intelligence, Shenzhen University of Advanced Technology
Jialong Li
Jialong Li
Waseda University
self-adaptive systemsrequirement engineeringhuman-in-the-loop