HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models Training

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of parallelizing Mixture-of-Experts (MoE) models across heterogeneous clusters, where architectural complexity and hardware disparities are difficult to reconcile simultaneously. We propose the first unified planning framework that jointly incorporates MoE awareness and cluster heterogeneity. The method constructs a lightweight cost model and employs a pruning-enhanced dynamic programming algorithm to efficiently search a six-dimensional parallelism space. It further supports non-uniform pipeline partitioning to accommodate complex hardware environments, generating training schedules directly deployable in Megatron-LM. Experimental results demonstrate up to a 3.2Γ— improvement in training throughput, with non-uniform partitioning contributing an additional 78% gain. The search completes in under one minute, highlighting both the efficiency and practical value of the proposed approach.
πŸ“ Abstract
As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training parallelism strategies at low cost while achieving superior performance. The difficulty of this problem is jointly determined by the complexity of the model and the underlying compute cluster. Meanwhile, mixture-of-experts (MoE) models are increasingly emerging as the dominant architecture and the rapid evolution of accelerator hardware has made cluster heterogeneity commonplace, posing substantial challenges to automatic parallelization. However, existing approaches typically target either MoE architectures or heterogeneous clusters, failing to generalize to scenarios where both challenges coexist. To this end, we present HAPMoE, a heterogeneity-aware automatic parallelism planner for MoE training. HAPMoE builds a lightweight MoE-aware cost model and efficiently searches a six-dimensional parallel space, producing parallel plans directly deployable on Megatron-LM. Experiments show that HAPMoE improves end-to-end training throughput by up to 3.2$\times$ over baselines across heterogeneous clusters. Its non-uniform pipeline partitioning yields an additional up to 78% gains, and its pruning-enhanced dynamic programming algorithm completes the search within 1 minute, demonstrating high efficiency and practical value in complex hardware environments.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Heterogeneous clusters
Automatic parallelism
Distributed training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Automatic Parallelism
Heterogeneous Clusters
Cost Model
Dynamic Programming
πŸ’Ό Related Jobs
No related jobs found.
M
Mengyuan Fan
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; Infinigence AI
P
Peizhuang Cong
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Zixiao Huang
Zixiao Huang
Phd student of Electronic Engineering, Tsinghua University
Machine Learning SystemDistributed SystemsDeep Learning Accelerator
S
Si Xu
Infinigence AI
Tong Qiao
Tong Qiao
Associate Professor, School of Cyberspace, Hangzhou Dianzi University
Media ForensicsAI SecurityDeepFake DetectionData Hiding
Yanghao Li
Yanghao Li
Apple
Computer Vision
J
Jing Yang
Infinigence AI
Tong Yang
Tong Yang
Peking University, Beijing, China. PKU. εŒ—δΊ¬ε€§ε­¦
SketchNetwork measurementBloom filterIP lookupHash Table
Q
Quanlu Zhang
Infinigence AI
Y
Yu Wang
Tsinghua University