Accurate Simulation of Distributed Training Jobs with Network Contention Modeling

📅 2026-09-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入MoSim模拟器,解决了现有分布式训练模拟器忽略或简单处理网络竞争问题,从而提高了模拟精度并减少了输入构建开销。
📝 Abstract
Trace-driven simulation is widely used to evaluate distributed training (DT) jobs in GPU clusters, but existing simulators either ignore network contention or approximate it with a fixed penalty. This misses how scheduling decisions determine which jobs share server network interfaces and inter-server links, thereby changing networking time during training. As a result, our motivating experiments demonstrate that they incur large errors, reaching up to 73.64% mean absolute percentage error (MAPE) in average job completion time (JCT). This paper introduces MoSim, a GPU-cluster simulator that models DT job execution under dynamic network contention. MoSim combines GPU-free characterization with network contention model: it obtains each job's compute time, networking time, and networking volume without GPUs, then uses the current worker assignment to estimate how shared network interfaces affect each job's iteration time. Our evaluation shows that, compared with existing simulators, MoSim reduces simulation error for average JCT by up to 3.28$\times$, tail (99th-percentile) JCT by up to 7.79$\times$, and makespan by up to 8.48$\times$, while modeling NIC contention factors with only 8.63% error on average. By avoiding real-GPU profiling, MoSim also reduces input construction overhead by 44.6$\times$.
Problem

Research questions and friction points this paper is trying to address.

Distributed Training
Network Contention
Simulation
GPU Clusters
Job Completion Time
Innovation

Methods, ideas, or system contributions that make the work stand out.

dynamic network contention
GPU-cluster simulator
MoSim
networking time estimation
shared network interfaces
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yeonho Yoo
Department of Computer Science and Artificial Intelligence, Dongguk University
H
Hyunho Lee
Department of Computer Science and Engineering, Korea University
H
Hyunmok Choi
Department of Computer Science and Engineering, Korea University
C
Chuck Yoo
Department of Computer Science and Engineering, Korea University
Gyeongsik Yang
Gyeongsik Yang
Korea University
Operating systemsNetwork virtualizationDatacenter networkingDistributed deep learning