Score
Designs and trains pretraining pipelines for neural architectures made of multiple expert subnetworks where a dynamic routing/gating mechanism selects which experts to activate per input, with the aim of inducing specialized, reusable experts and capturing multi-task priors. This work involves building expert parameterizations and dynamic routing policies, specifying losses and regularizers that encourage diversity and reuse, and analyzing expert activation patterns, specialization, efficiency, and transfer performance.
This work systematically dissects the multidimensional design space of Mixture-of-Experts (MoE) architectures in large language models, moving beyond conventional generational narratives. It introduces a five-dimensional analytical framework encompassing expert granularity, topology, routing flexibility, load-balancing scope, and execution structure, and constructs a dependency graph to elucidate the coupling mechanisms across four control planes: expert topology, routing, load balancing, and expert parallelism. The framework’s validity is empirically demonstrated through iso-budget pretraining experiments integrating algorithmic innovations—such as Top-k routing, shared and fine-grained experts, and dynamic expert composition—with system-level optimizations including token dispatch, device placement, and all-to-all communication. The study further distills key open challenges for the future development of MoE systems.
Dynamic routing in Mixture-of-Experts (MoE) models is vulnerable to redundant neuron activations, leading to biased expert selection and insufficient expert diversity. Method: We propose a neural suppression–enhanced dynamic routing mechanism that applies learnable suppression signals to redundant neuron populations within the shared feature space prior to routing decisions, explicitly attenuating collinear responses to improve discriminability and specialization of expert path selection. Contribution/Results: Unlike prior MoE approaches, this work is the first to systematically demonstrate the routing-quality benefits of neural suppression and integrate it end-to-end into Transformer-like architectures without increasing parameter count. Experiments across multiple NLP and vision benchmarks show consistent improvements: +1.2–2.8% accuracy gains, −37% reduction in routing variance, and enhanced expert utilization balance and task adaptability—establishing a novel paradigm for efficient, diverse sparse modeling.
This work proposes DynaMoE, a novel Mixture-of-Experts (MoE) framework that overcomes the limitations of conventional MoE architectures, which rely on fixed Top-K routing and uniform expert allocation, thereby struggling to adapt to varying input complexity and task characteristics. DynaMoE introduces a dynamic token-level routing mechanism alongside six inter-layer expert capacity scheduling strategies—such as decreasing, pyramid, and wave patterns—to break free from static constraints on expert count and assignment. Theoretical analysis grounded in gradient variance and computational efficiency substantiates the approach’s efficacy. Experiments demonstrate that DynaMoE significantly outperforms static baselines on both image classification and language modeling tasks, with the optimal scheduling strategy depending on task type and model scale. Moreover, the method effectively reduces training gradient variance, enhancing model expressivity and stability.
This work addresses the lack of convergence guarantees for soft-routing Mixture-of-Experts (MoE) models under joint training of nonlinear routers and experts. We propose a provably correct feature learning framework grounded in a student–teacher paradigm. Methodologically, we model a moderately overparameterized MoE architecture, incorporate dynamic weighted aggregation and soft routing, and design a pruning-augmented fine-tuning strategy with provable convergence. Theoretically, we establish the first global convergence guarantee for the student network under joint training, proving exact recovery of teacher parameters. We further uncover an intrinsic gradient-guided mechanism by which experts shape router learning. Finally, we deliver a practical yet theoretically sound optimization paradigm: pruning preserves performance, and fine-tuning enjoys rigorous convergence guarantees. This work provides the first unified, interpretable theoretical lens into MoE training dynamics.
This work investigates the opaque expert specialization mechanism in Mixture-of-Experts (MoE) models, which limits inference and memory efficiency. By analyzing domain-specific routing patterns and employing an early-decoding framework, the study systematically examines how individual experts contribute to model outputs. Through comprehensive analyses—including routing distribution statistics, cosine similarity of hidden states, comparisons between single-expert and ensemble outputs, and perplexity evaluation—the authors find that a small subset of experts handles over 50% of all requests. Remarkably, outputs from a single dominant expert exhibit high consistency with the full model (cosine similarity up to 0.95), with only a 5% increase in perplexity. These findings suggest that precise expert pruning can substantially enhance inference efficiency without compromising performance, offering a promising avenue for efficient MoE deployment and knowledge localization.
本文通过局部聚合视角分析混合专家模型,探讨路由、稀疏激活和共享专家等设计选择的统计作用,分离出逼近误差、专家学习误差和路由器估计误差。
This work addresses the pervasive issue of deep routing collapse in large Mixture-of-Experts (MoE) models for low-resource languages, which leads to imbalanced expert utilization and constrained multilingual capabilities. The study reveals, for the first time, that this phenomenon stems from insufficient pretraining data rather than inherent linguistic properties. To diagnose multilingual capacity, the authors propose routing entropy and expert specialization as key indicators. Through balanced bilingual continual pretraining (CPT) and supervised fine-tuning (SFT) on Hebrew, Japanese, and other languages using both pure Transformer and Mamba-Transformer hybrid architectures, they demonstrate that CPT substantially increases routing entropy, encourages language-agnostic expert sharing, and consistently enhances downstream performance, whereas SFT yields limited gains. These findings underscore the critical role of data balance in scaling MoE models multilingually.
This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.
This work systematically investigates the interplay among key design dimensions in Mixture-of-Experts (MoE) architectures—such as the number of experts, expert granularity, heterogeneity, shared experts, and load balancing—through over 2,000 large-scale pretraining experiments. The study reveals that the number of experts and their granularity are the dominant factors governing model performance, while other design choices exert comparatively limited influence. Notably, increasing the total MoE parameters consistently enhances performance across all active parameter budgets, and the optimal expert size is determined solely by the number of active parameters. Furthermore, the effectiveness of dropless routing is empirically validated, demonstrating consistent performance gains.
This work addresses the routing collapse and expert deadlocks that commonly afflict Token-Choice sparse Mixture-of-Experts (MoE) architectures in video diffusion Transformers, which severely limit expert diversity utilization. Starting from a 5-billion-parameter dense model, the authors formulate three principles for converting dense networks to MoE. Through temporal routing analysis of 65 million tokens, they reveal that deadlocked layers follow a U-shaped distribution across the network depth and propose a “functional redundancy” hypothesis to explain this phenomenon. Building on these insights, they integrate expert cloning, zero-initialized gating, auxiliary losses, and enhanced router designs—including linear, MLP, and cross-attention variants—to effectively mitigate bfloat16 precision pitfalls. Their approach alleviates single-expert deadlocks in approximately two-thirds of network layers, endows the model with partial self-recovery capability, delineates the capacity limits of the Token-Choice paradigm, and outlines a three-stage roadmap toward unified vision models and ultimately world models.