moe long-context scaling

Designs, builds, and analyzes training and inference systems for Mixture‑of‑Experts (MoE) models that scale input context windows to extremely long lengths (up to million‑token sequences), including expert routing and capacity mechanisms, activation and sequence‑memory management, and memory‑saving/offloading techniques so training and pretraining proceed losslessly and without out‑of‑memory failures.

moelong-contextscaling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.49
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Mixture of Experts in Large Language Models

Jul 15, 2025
DZ
Danyang Zhang
🏛️ ByteDance Inc | Imperial College London | Purdue University | Heriot-Watt University | Vokram Group | Singapore General Hospital

This paper presents a systematic survey of recent advances in Mixture-of-Experts (MoE) architectures for large language models. Addressing the fundamental trade-off between model capacity scaling and computational efficiency, it investigates key directions: expert gating and dynamic routing mechanisms, hierarchical sparse structure design, meta-learning–enhanced expert collaboration, multimodal/multitask adaptation, and practical deployment challenges. The work proposes a novel MoE effectiveness enhancement framework centered on expert diversity modeling, gating calibration optimization, and improved reliability of inference-time expert aggregation—demonstrating significant gains over both dense models and Bayesian baselines of comparable parameter count. Beyond empirical advances, the study identifies critical bottlenecks—including expert load imbalance, training instability, and hardware inefficiency—and establishes a principled theoretical framework alongside actionable guidelines for designing efficient, scalable MoE-based LLMs. (149 words)

Addressing challenges in expert diversity and inference reliabilityAnalyzing expert gating and routing mechanisms in MoEEnhancing model performance with minimal computational overhead

To address GPU memory exhaustion and PCIe transfer latency caused by expert prefetching failures in MoE model inference, this paper proposes an acceleration method leveraging expert redundancy. The core innovation is the first introduction of a dynamic functional-substitution mechanism: when the target expert misses the cache, a semantically similar expert—already loaded—is dynamically scheduled for inference, eliminating stalls or performance degradation. The method comprises expert similarity modeling, runtime redundant scheduling, CPU-GPU collaborative execution, lightweight prefetch prediction, and cache-aware expert placement. Experiments under memory-constrained settings demonstrate that our approach reduces end-to-end latency by up to 42%, improves throughput by up to 3.1×, and maintains accuracy loss below 0.3%, closely approaching the performance of full-expert loading.

Accelerating MoE inference under GPU memory constraintsMaintaining model accuracy when prefetching mechanisms failReducing latency from expert offloading across PCIe interconnect

A Closer Look into Mixture-of-Experts in Large Language Models

Jun 26, 2024
KM
Ka Man Lo
🏛️ University of Macau | University of Edinburgh | Tsinghua University | INF Technology | HKUST

The internal mechanisms and modular nature of Mixture-of-Experts (MoE) large language models remain poorly understood, particularly regarding expert granularity, routing behavior, and layer-wise expert diversity. Method: We conduct attribution analysis, expert activation visualization, output norm statistics, and controlled experiments across three representative MoE architectures—Mixtral, GLaM, and DeepSpeed-MoE. Contribution/Results: We empirically establish that individual neurons function as fine-grained experts; routers exhibit strong preference for high-norm experts; and expert diversity generally increases with network depth—except in the final layer. Based on these findings, we formulate a hierarchical evolution law of expert diversity and provide actionable guidelines for router design and expert allocation. Our work formally validates the modular architecture of MoE models, identifies anomalous behavior in the top layer, and has directly informed routing strategy improvements across multiple research teams. The open-sourced code has garnered significant community attention.

Exploring parametric and behavioral features of MoE modelsInvestigating expert diversity and router selection mechanismsUnderstanding inner workings of MoE-based large language models

This work investigates the opaque expert specialization mechanism in Mixture-of-Experts (MoE) models, which limits inference and memory efficiency. By analyzing domain-specific routing patterns and employing an early-decoding framework, the study systematically examines how individual experts contribute to model outputs. Through comprehensive analyses—including routing distribution statistics, cosine similarity of hidden states, comparisons between single-expert and ensemble outputs, and perplexity evaluation—the authors find that a small subset of experts handles over 50% of all requests. Remarkably, outputs from a single dominant expert exhibit high consistency with the full model (cosine similarity up to 0.95), with only a 5% increase in perplexity. These findings suggest that precise expert pruning can substantially enhance inference efficiency without compromising performance, offering a promising avenue for efficient MoE deployment and knowledge localization.

expert specializationinference optimizationMixture of Experts

Latest Papers

What's happening recently
View more

This work addresses the memory bottlenecks and communication overheads encountered when training trillion-parameter Mixture-of-Experts (MoE) models with million-token context lengths. To overcome these challenges, the authors propose a “Mixture-of-Parallelisms” paradigm that synergistically integrates data, tensor, expert, and pipeline parallelism, complemented by memory-efficient optimizer state management and communication scheduling strategies. This approach enables, for the first time, lossless training of trillion-parameter MoE models at 1M-token context lengths while substantially reducing hardware requirements. Experimental results demonstrate that on a cluster of 12 nodes equipped with 8×H200 GPUs each, the method achieves per-GPU throughput 4.7–8.2× higher than the FSDP2 baseline, which suffers from out-of-memory errors even at context lengths of 64–128K tokens.

large-scale traininglong context lengthmemory efficiency

This work addresses the coupled memory, communication, and computation bottlenecks in large-scale Mixture-of-Experts (MoE) model training by proposing a full-stack co-optimization framework. Central to this approach is the Parallel Folding multi-dimensional parallelism strategy, which integrates fine-grained recomputation, expert scheduling and offloading, Grouped GEMM kernel fusion, CUDA Graphs, and FP8/NVFP4 low-precision training to enable highly efficient overlap of communication and computation. The framework supports scalable training of MoE models ranging from billions to trillions of parameters across thousands of GPUs. On NVIDIA GB300/GB200 systems, it achieves 1,233/1,048 TFLOPS/GPU for DeepSeek-V3-685B and 974/919 TFLOPS/GPU for Qwen3-235B, significantly advancing the system efficiency and accessibility of large-scale MoE training.

large-scale MoEmemory-computation-communication trade-offMixture-of-Experts

This work addresses the substantial cross-node communication overhead in multi-node Mixture-of-Experts (MoE) inference, which stems from imbalanced expert loads and inefficient token routing. For the first time, it systematically characterizes three key properties of MoE expert activation: dynamic load imbalance, task-domain-dependent expert preferences, and strong correlation between the prefill and decode phases. Leveraging these insights, the authors propose a workload-aware microbatch grouping and expert placement strategy that enhances token-expert locality. Evaluated on over 100,000 real-world activation traces across multiple MoE models and datasets, the approach reduces all-to-all communication volume by up to 20×, significantly lowering decoding latency and improving accelerator utilization.

inter-node communicationload imbalanceMixture-of-Experts

This work systematically investigates the interplay among key design dimensions in Mixture-of-Experts (MoE) architectures—such as the number of experts, expert granularity, heterogeneity, shared experts, and load balancing—through over 2,000 large-scale pretraining experiments. The study reveals that the number of experts and their granularity are the dominant factors governing model performance, while other design choices exert comparatively limited influence. Notably, increasing the total MoE parameters consistently enhances performance across all active parameter budgets, and the optimal expert size is determined solely by the number of active parameters. Furthermore, the effectiveness of dropless routing is empirically validated, demonstrating consistent performance gains.

expert countexpert granularityload balancing

Hot Scholars

MR

Mehdi Rezagholizadeh

Principal Research Scientist, Advanced Micro Devices (AMD)
Efficient AINLP/Computer VisionDeep Learning
XZ

Xiangyu Zhang

Co-founder & Chief Scientist of StepFun
Neural Network ArchitecturesEfficient Deep LearningComputer Vision
PF

Parsa Farinneya

University of Toronto
Machine learningNLPRecommender system
AS

Aman Sharma

PhD Student, KTH Royal Institute of Technology
Software EngineeringSoftware Supply Chain