layer-adaptive moe

Designs and implements neural mixture-of-experts architectures that allocate, share, and adapt experts at the level of individual network layers, including layer-wise shared experts and routing mechanisms. Builds and analyzes dynamic expert-selection and usage-update mechanisms (e.g., momentum-based updates) to maintain cross-task shared representations and promote transfer of knowledge across related tasks.

layer-adaptivemoe

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This work addresses a critical gap in existing literature by providing the first comprehensive survey of sparse Mixture-of-Experts (MoE) models, systematically integrating their algorithmic foundations, decentralized architectures, and applications in vertical domains. It thoroughly examines core mechanisms such as routing strategies and expert network design, while further extending the discussion to decentralized deployment paradigms and adaptation methods for cross-modal and domain-specific scenarios. By synthesizing recent advances across these dimensions, this survey fills a notable void in the current body of review literature and offers an authoritative reference for researchers and practitioners aiming to develop efficient, scalable large models grounded in sparse MoE principles.

Decentralized ArchitecturesRouting NetworkSparse Mixture-of-Experts

Must-Read Papers

Most classic and influential ideas
View more

This work systematically dissects the multidimensional design space of Mixture-of-Experts (MoE) architectures in large language models, moving beyond conventional generational narratives. It introduces a five-dimensional analytical framework encompassing expert granularity, topology, routing flexibility, load-balancing scope, and execution structure, and constructs a dependency graph to elucidate the coupling mechanisms across four control planes: expert topology, routing, load balancing, and expert parallelism. The framework’s validity is empirically demonstrated through iso-budget pretraining experiments integrating algorithmic innovations—such as Top-k routing, shared and fine-grained experts, and dynamic expert composition—with system-level optimizations including token dispatch, device placement, and all-to-all communication. The study further distills key open challenges for the future development of MoE systems.

expert parallelismexpert topologyload balancing

Load Balancing Mixture of Experts with Similarity Preserving Routers

Jun 16, 2025
NO
Nabil Omi
🏛️ University of Washington | Microsoft Research | Allen Institute for AI

Sparse Mixture-of-Experts (MoE) models often suffer from capacity waste and performance degradation due to routing bias toward a small subset of experts. Conventional load-balancing methods enforce uniform expert utilization but risk undermining semantic coherence, leading to knowledge redundancy across experts. To address this, we propose a similarity-preserving load-balancing mechanism: a differentiable routing loss grounded in token embedding similarity, which encourages semantically similar tokens to be consistently routed to the same expert—thereby jointly optimizing load distribution and routing consistency. Our approach requires no additional experts or auxiliary modules and integrates seamlessly into standard MoE training pipelines. Experiments demonstrate a 36% acceleration in convergence on benchmark tasks, substantial reduction in inter-expert knowledge redundancy, and improved model generalization and inference efficiency.

Ensures consistent expert assignment for similar inputsPrevents expert underutilization in sparse MoE modelsReduces redundant knowledge learning in expert routing

本文提出MoRE模型,通过在相邻层间共享专家池来解决Mixture-of-Experts架构中参数增加导致内存占用高的问题,同时引入深度嵌入以区分不同层。

language modelingmemory footprintMixture-of-Experts

This study investigates the intrinsic mechanisms underlying multilingual capabilities in Mixture-of-Experts (MoE) large language models, with a focus on cross-lingual differences in routing behavior, expert specialization, and layer-wise processing. Through systematic analysis of routing strategies and expert activation patterns, the work reveals for the first time that high-resource languages tend to share experts, whereas low-resource languages prefer dedicated experts, with intermediate layers functioning as language-agnostic capacity hubs. Building on these insights, the authors propose a hierarchical routing guidance method during inference that dynamically steers typologically related languages toward shared experts. Experimental results demonstrate consistent performance gains across multilingual tasks, with particularly notable improvements for language pairs within the same linguistic family.

expert specializationlayerwise steeringMixture-of-Experts

This work investigates the opaque expert specialization mechanism in Mixture-of-Experts (MoE) models, which limits inference and memory efficiency. By analyzing domain-specific routing patterns and employing an early-decoding framework, the study systematically examines how individual experts contribute to model outputs. Through comprehensive analyses—including routing distribution statistics, cosine similarity of hidden states, comparisons between single-expert and ensemble outputs, and perplexity evaluation—the authors find that a small subset of experts handles over 50% of all requests. Remarkably, outputs from a single dominant expert exhibit high consistency with the full model (cosine similarity up to 0.95), with only a 5% increase in perplexity. These findings suggest that precise expert pruning can substantially enhance inference efficiency without compromising performance, offering a promising avenue for efficient MoE deployment and knowledge localization.

expert specializationinference optimizationMixture of Experts

Latest Papers

What's happening recently
View more

This work systematically investigates the interplay among key design dimensions in Mixture-of-Experts (MoE) architectures—such as the number of experts, expert granularity, heterogeneity, shared experts, and load balancing—through over 2,000 large-scale pretraining experiments. The study reveals that the number of experts and their granularity are the dominant factors governing model performance, while other design choices exert comparatively limited influence. Notably, increasing the total MoE parameters consistently enhances performance across all active parameter budgets, and the optimal expert size is determined solely by the number of active parameters. Furthermore, the effectiveness of dropless routing is empirically validated, demonstrating consistent performance gains.

expert countexpert granularityload balancing

This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.

dynamic routingexpert capacityMixture-of-Experts

This work addresses the pervasive issue of deep routing collapse in large Mixture-of-Experts (MoE) models for low-resource languages, which leads to imbalanced expert utilization and constrained multilingual capabilities. The study reveals, for the first time, that this phenomenon stems from insufficient pretraining data rather than inherent linguistic properties. To diagnose multilingual capacity, the authors propose routing entropy and expert specialization as key indicators. Through balanced bilingual continual pretraining (CPT) and supervised fine-tuning (SFT) on Hebrew, Japanese, and other languages using both pure Transformer and Mamba-Transformer hybrid architectures, they demonstrate that CPT substantially increases routing entropy, encourages language-agnostic expert sharing, and consistently enhances downstream performance, whereas SFT yields limited gains. These findings underscore the critical role of data balance in scaling MoE models multilingually.

expert entropylow-resource languagesMixture-of-Experts

Hot Scholars

MV

Marcos Villagra

Bagel Labs 🥯
Computational ComplexityMachine LearningCryptographyPrivacy
ZJ

Zhiying Jiang

University of Waterloo
Natural Language ProcessingMachine Learning
ZL

Zhixuan Liang

University of Hong Kong
Embodied AIMachine LearningRoboticsComputer Vision
XG

Xinping Guan

Shanghai Jiao Tong University
Wireless Networks and ApplicationsInternet of ThingsControl and Systems