Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts

📅 2024-11-02
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address communication overhead, computational redundancy, and excessive memory consumption in training Mixture-of-Experts (MoE) models on heterogeneous hardware, this paper proposes the first hardware-aware expert assignment framework. Our method introduces three key innovations: (1) expert-specific operators enabling zero-redundancy in-place computation; (2) a dual-centered (data- and model-driven) adaptive parallelism configuration mechanism; and (3) device-level pipelined shared caching. Evaluated under realistic heterogeneous cluster settings, the framework preserves model accuracy while reducing memory footprint by 10–48% and accelerating training throughput by 0.5–4.3×. Consequently, end-to-end training latency is significantly lowered. This work provides a systematic solution for efficient large-scale MoE deployment across heterogeneous infrastructure.

Technology Category

Machine Learning: Mixture of Experts (MoE)Search and Optimization: Mixed Discrete/Continuous SearchData Mining & Knowledge Management: Scalability, Parallel & Distributed Systems

Application Category

User Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Mixture-of-Experts (MoE) has emerged as a practical approach to scale up parameters for the Transformer model to achieve better generalization while maintaining a sub-linear increase in computation overhead. Current MoE models are mainly built with expert parallelism on distributed devices. However, it usually depends on homogeneous devices to deploy and suffers from heavy communication overhead and computation redundancy. In this paper, we explore developing a exttt{H}eterogeneous-aware exttt{EX}pert exttt{A}llocation framework, extbf{ exttt{HEXA-MoE}}, with significantly enhanced computing efficiency. It contains two components: ($1$) extit{Expert-Specific Operators}. We replace the typical general matrix multiplication or grouped matrix multiplication interfaces with our operators, which allows the computing to be performed in an in-place manner with extbf{ZERO} redundancy. ($2$) extit{Adaptive Data- and Model-Centric Configurations} for different workload scales. Specifically, we introduce a pipeline-shared cache on each device to tackle the heavy memory consumption in the existing data-centric MoE library. Comprehensive experiments on the Swin-MoE benchmark consistently reveal the effectiveness of our exttt{HEXA-MoE} framework, i.e., reducing $10%sim48%$ memory consumption and achieving $0.5sim4.3 imes$ speed up compared to current state-of-the-art MoE libraries. Furthermore, we examine our exttt{HEXA-MoE} with heterogeneous devices for both data- and model-centric settings. Promising results show that employing optimal parallel configuration with exttt{HEXA-MoE} on heterogeneous devices can substantially minimize overall latency. Codes are available at https://github.com/UNITES-Lab/HEXA-MoE.
Problem

Research questions and friction points this paper is trying to address.

Reducing communication overhead in MoE models
Optimizing computation redundancy in heterogeneous devices
Improving memory efficiency in data-centric MoE libraries
Innovation

Methods, ideas, or system contributions that make the work stand out.

Heterogeneous-aware expert allocation framework
Expert-Specific Operators for zero redundancy
Adaptive data- and model-centric configurations
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Peking University | University of Science and Technology of China | University of North Carolina