HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models Training
This study addresses the challenge of parallelizing Mixture-of-Experts (MoE) models across heterogeneous clusters, where architectural complexity and hardware disparities are difficult to reconcile simultaneously. We propose the first unified planning framework that jointly incorporates MoE awareness and cluster heterogeneity. The method constructs a lightweight cost model and employs a pruning-enhanced dynamic programming algorithm to efficiently search a six-dimensional parallelism space. It further supports non-uniform pipeline partitioning to accommodate complex hardware environments, generating training schedules directly deployable in Megatron-LM. Experimental results demonstrate up to a 3.2× improvement in training throughput, with non-uniform partitioning contributing an additional 78% gain. The search completes in under one minute, highlighting both the efficiency and practical value of the proposed approach.