🤖 AI Summary
To address communication overhead, computational redundancy, and excessive memory consumption in training Mixture-of-Experts (MoE) models on heterogeneous hardware, this paper proposes the first hardware-aware expert assignment framework. Our method introduces three key innovations: (1) expert-specific operators enabling zero-redundancy in-place computation; (2) a dual-centered (data- and model-driven) adaptive parallelism configuration mechanism; and (3) device-level pipelined shared caching. Evaluated under realistic heterogeneous cluster settings, the framework preserves model accuracy while reducing memory footprint by 10–48% and accelerating training throughput by 0.5–4.3×. Consequently, end-to-end training latency is significantly lowered. This work provides a systematic solution for efficient large-scale MoE deployment across heterogeneous infrastructure.
📝 Abstract
Mixture-of-Experts (MoE) has emerged as a practical approach to scale up parameters for the Transformer model to achieve better generalization while maintaining a sub-linear increase in computation overhead. Current MoE models are mainly built with expert parallelism on distributed devices. However, it usually depends on homogeneous devices to deploy and suffers from heavy communication overhead and computation redundancy. In this paper, we explore developing a exttt{H}eterogeneous-aware exttt{EX}pert exttt{A}llocation framework, extbf{ exttt{HEXA-MoE}}, with significantly enhanced computing efficiency. It contains two components: ($1$) extit{Expert-Specific Operators}. We replace the typical general matrix multiplication or grouped matrix multiplication interfaces with our operators, which allows the computing to be performed in an in-place manner with extbf{ZERO} redundancy. ($2$) extit{Adaptive Data- and Model-Centric Configurations} for different workload scales. Specifically, we introduce a pipeline-shared cache on each device to tackle the heavy memory consumption in the existing data-centric MoE library. Comprehensive experiments on the Swin-MoE benchmark consistently reveal the effectiveness of our exttt{HEXA-MoE} framework, i.e., reducing $10%sim48%$ memory consumption and achieving $0.5sim4.3 imes$ speed up compared to current state-of-the-art MoE libraries. Furthermore, we examine our exttt{HEXA-MoE} with heterogeneous devices for both data- and model-centric settings. Promising results show that employing optimal parallel configuration with exttt{HEXA-MoE} on heterogeneous devices can substantially minimize overall latency. Codes are available at https://github.com/UNITES-Lab/HEXA-MoE.