π€ AI Summary
This study addresses the challenge of parallelizing Mixture-of-Experts (MoE) models across heterogeneous clusters, where architectural complexity and hardware disparities are difficult to reconcile simultaneously. We propose the first unified planning framework that jointly incorporates MoE awareness and cluster heterogeneity. The method constructs a lightweight cost model and employs a pruning-enhanced dynamic programming algorithm to efficiently search a six-dimensional parallelism space. It further supports non-uniform pipeline partitioning to accommodate complex hardware environments, generating training schedules directly deployable in Megatron-LM. Experimental results demonstrate up to a 3.2Γ improvement in training throughput, with non-uniform partitioning contributing an additional 78% gain. The search completes in under one minute, highlighting both the efficiency and practical value of the proposed approach.
π Abstract
As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training parallelism strategies at low cost while achieving superior performance. The difficulty of this problem is jointly determined by the complexity of the model and the underlying compute cluster. Meanwhile, mixture-of-experts (MoE) models are increasingly emerging as the dominant architecture and the rapid evolution of accelerator hardware has made cluster heterogeneity commonplace, posing substantial challenges to automatic parallelization. However, existing approaches typically target either MoE architectures or heterogeneous clusters, failing to generalize to scenarios where both challenges coexist. To this end, we present HAPMoE, a heterogeneity-aware automatic parallelism planner for MoE training. HAPMoE builds a lightweight MoE-aware cost model and efficiently searches a six-dimensional parallel space, producing parallel plans directly deployable on Megatron-LM. Experiments show that HAPMoE improves end-to-end training throughput by up to 3.2$\times$ over baselines across heterogeneous clusters. Its non-uniform pipeline partitioning yields an additional up to 78% gains, and its pruning-enhanced dynamic programming algorithm completes the search within 1 minute, demonstrating high efficiency and practical value in complex hardware environments.