Score
Designs and implements neural architectures and routing mechanisms that apply sparse mixture-of-experts over the time and frequency axes of a time–frequency representation, e.g., selecting small sets of experts per time frame or frequency band and alternating time-wise and frequency-wise expert groups. Builds and evaluates expert modules, sparse routing/scheduling algorithms, and compute–accuracy trade-offs to increase model capacity while keeping runtime and memory cost low.
This work addresses a critical gap in existing literature by providing the first comprehensive survey of sparse Mixture-of-Experts (MoE) models, systematically integrating their algorithmic foundations, decentralized architectures, and applications in vertical domains. It thoroughly examines core mechanisms such as routing strategies and expert network design, while further extending the discussion to decentralized deployment paradigms and adaptation methods for cross-modal and domain-specific scenarios. By synthesizing recent advances across these dimensions, this survey fills a notable void in the current body of review literature and offers an authoritative reference for researchers and practitioners aiming to develop efficient, scalable large models grounded in sparse MoE principles.
本文通过局部聚合视角分析混合专家模型,探讨路由、稀疏激活和共享专家等设计选择的统计作用,分离出逼近误差、专家学习误差和路由器估计误差。
This work addresses the limitations of conventional sparse mixture-of-experts (MoE) models, which employ independent routing at each layer, resulting in an excessively large path space and poor statistical efficiency that hinder the learning of stable expert routing structures. To overcome this, the authors propose Path-Constrained Mixture of Experts (PathMoE), a novel architecture that shares router parameters across layers to dramatically reduce the effective path space, thereby enhancing path consistency and structural learnability. Notably, PathMoE naturally induces token clustering according to linguistic functionality without requiring auxiliary load-balancing losses. Experiments demonstrate that PathMoE achieves lower perplexity, superior downstream task performance, and greater robustness to routing perturbations compared to standard MoE baselines, consistently across both 0.9B and 16B parameter scales.
This study investigates the impact of expert granularity on model performance within dense Mixture-of-Experts (MoE) architectures under a fixed computational budget. Theoretically, we demonstrate that dense MoE is equivalent to a standard feed-forward network (FFN) with token-dependent scaling, revealing a non-monotonic effect of fine-grained experts. Empirically, employing SwiGLU activations, Softmax gating, and fixed-width FFNs, we systematically compare varying numbers of activated experts in terms of validation loss and routing behavior. Our findings indicate that activating K=2 experts yields optimal performance, significantly outperforming the baseline, whereas higher granularity paradoxically degrades results. These observations confirm the critical role of soft gating mechanisms under specific architectural configurations.
In sparsely activated Mixture-of-Experts (MoE) models, conventional top-k routing enforces fixed expert capacity constraints, leading to token dropping or inefficient padding and thus degrading hardware utilization; removing such constraints, however, causes severe load imbalance and reduced computational efficiency. To address this, we propose MaxScore routing—a novel mechanism that formulates token assignment as a minimum-cost maximum-flow problem. By integrating a differentiable SoftTopk operator with graph-based flow optimization, MaxScore achieves dynamic load balancing without explicit capacity limits, avoiding the limitations of iterative rerouting and optimal transport while preserving both differentiability and global optimality. Experiments demonstrate that, at identical FLOPs, models trained with MaxScore achieve lower training loss and higher downstream task performance, significantly improving computational efficiency and overall model effectiveness.
This work systematically dissects the multidimensional design space of Mixture-of-Experts (MoE) architectures in large language models, moving beyond conventional generational narratives. It introduces a five-dimensional analytical framework encompassing expert granularity, topology, routing flexibility, load-balancing scope, and execution structure, and constructs a dependency graph to elucidate the coupling mechanisms across four control planes: expert topology, routing, load balancing, and expert parallelism. The framework’s validity is empirically demonstrated through iso-budget pretraining experiments integrating algorithmic innovations—such as Top-k routing, shared and fine-grained experts, and dynamic expert composition—with system-level optimizations including token dispatch, device placement, and all-to-all communication. The study further distills key open challenges for the future development of MoE systems.
This work investigates the information-theoretic efficiency of routing mechanisms in sparse Mixture-of-Experts (MoE) architectures, aiming to balance model accuracy with communication and computational resource utilization. The gating router is modeled as a stochastic channel, and a discrete mutual information estimator is proposed under a finite expert pool. Empirical posterior distributions \( q(W|S) \) are leveraged to compute \( I(X;T) \) and \( I(S;W) \), with the latter shown to exhibit a monotonic relationship with the generalization gap. The Blahut–Arimoto algorithm is employed to trace the accuracy–rate trade-off curve. Experiments demonstrate that the proposed mutual information estimator effectively tracks the generalization gap and significantly outperforms both the Xu–Raginsky bound and the uniform joint bound, offering a practical analytical tool for resource-aware MoE systems.
This work addresses the instability and optimization challenges commonly encountered in training sparse vision Mixture-of-Experts (MoE) models, which often stem from gradient blocking and insufficient feedback during routing. To mitigate these issues, the authors propose a Teacher-Guided Routing mechanism (TGR-MoE), which, for the first time, leverages intermediate representations from a pretrained dense teacher model to provide knowledge-driven pseudo-supervisory signals for the student router, thereby enabling stable routing decisions early in training. By integrating knowledge distillation, sparse MoE architecture, and a representation-similarity-based routing strategy, TGR-MoE achieves substantial improvements in both accuracy and routing consistency on ImageNet-1K and CIFAR-100, maintaining robust training stability even under highly sparse configurations.
This work addresses the fundamental trade-off in sparse Mixture-of-Experts (MoE) models between load balancing and expert specialization, which often leads to routing collapse or diminished expert diversity. The authors propose Hi-MoE, a novel framework that decomposes routing into two coupled hierarchical levels: inter-group routing ensures balanced token distribution across expert groups, while intra-group routing fosters complementary expert specialization and prevents collapse. This principled redesign of router behavior consistently outperforms existing sparse routing and grouped MoE approaches across both NLP and vision benchmarks. In a 58B-token pretraining setting, Hi-MoE-7B achieves a 5.6% lower perplexity and 40% improved expert balance compared to OLMoE-7B.
It remains unclear whether the routing mechanism in sparse Mixture-of-Experts (MoE) models exhibits task-conditioned behavior. This work proposes the concept of “routing signatures” to characterize the activation patterns of experts across layers for different tasks and provides the first empirical evidence that MoE routing is highly task-sensitive. By constructing routing signatures, defining similarity metrics, applying logistic regression classification, and analyzing inter-layer routing signals, the study reveals that routing signatures of tasks within the same category exhibit a similarity of 0.8435, significantly higher than the 0.6225 observed across categories (Cohen’s d = 1.44). Remarkably, routing signatures alone achieve 92.5% accuracy in four-way task classification. The authors release the MOE-XRAY toolkit to advance interpretability research in MoE models, establishing a new paradigm for analyzing expert routing dynamics.
This work proposes a novel mixture-of-experts (MoE) architecture that eliminates the need for explicit routing mechanisms commonly found in traditional MoE models. By embedding activation logic directly within each expert and enabling end-to-end continuous gradient flow, experts autonomously determine their own activation without reliance on external routers, Softmax operations, Top-K selection, or hard-coded load-balancing heuristics. The approach introduces a unified, adaptive load-balancing framework that jointly optimizes resource allocation across both experts and tokens, supporting configurable dual-objective balancing. Experimental results demonstrate that the proposed model consistently outperforms existing baselines across multiple benchmarks, exhibiting superior scalability and robustness while removing rigid inductive biases imposed by centralized routing.