Score
Designs and implements modules and training procedures that dynamically add, grow, prune, and initialize expert submodules within a model to adjust representational capacity without retraining the entire system. This work includes algorithms for deciding when to expand or prune experts, methods for rank growth with orthogonal initialization, and strategies to scale the expert set while maintaining a fixed parameter budget.
This paper addresses two critical challenges in large language model development: excessive computational overhead and difficulty in modeling heterogeneous, complex data. To tackle these, we present a systematic, up-to-date survey of Mixture-of-Experts (MoE) models. Unlike prior surveys—often outdated or narrowly scoped—we unify and analyze MoE advancements across emerging paradigms including continual learning, meta-learning, and reinforcement learning. We propose a comprehensive framework integrating theoretical analysis (e.g., convergence guarantees), multimodal adaptation (vision and language), and systems-level optimizations (sparse routing, load balancing, distributed training). Furthermore, we introduce a taxonomy of future research directions. Our work establishes the most complete MoE knowledge graph to date, explicitly identifying key bottlenecks and viable technical pathways. It serves as both a methodological foundation and an engineering roadmap for developing efficient, scalable large models.
This work addresses catastrophic forgetting in fine-tuning pretrained models, where newly acquired knowledge overwrites previously learned information. To mitigate this issue, the authors propose a function-preserving model expansion approach that mathematically duplicates and scales parameters of selected Transformer submodules during initialization. This technique enables stable training and faithful retention of original model capabilities without altering the initial functionality. By circumventing the traditional trade-off between plasticity and stability, the method achieves performance comparable to full fine-tuning while expanding only a minimal number of layers. Consequently, it fully preserves the model’s original knowledge and substantially reduces computational overhead.
This work systematically dissects the multidimensional design space of Mixture-of-Experts (MoE) architectures in large language models, moving beyond conventional generational narratives. It introduces a five-dimensional analytical framework encompassing expert granularity, topology, routing flexibility, load-balancing scope, and execution structure, and constructs a dependency graph to elucidate the coupling mechanisms across four control planes: expert topology, routing, load balancing, and expert parallelism. The framework’s validity is empirically demonstrated through iso-budget pretraining experiments integrating algorithmic innovations—such as Top-k routing, shared and fine-grained experts, and dynamic expert composition—with system-level optimizations including token dispatch, device placement, and all-to-all communication. The study further distills key open challenges for the future development of MoE systems.
This work systematically investigates the interplay among key design dimensions in Mixture-of-Experts (MoE) architectures—such as the number of experts, expert granularity, heterogeneity, shared experts, and load balancing—through over 2,000 large-scale pretraining experiments. The study reveals that the number of experts and their granularity are the dominant factors governing model performance, while other design choices exert comparatively limited influence. Notably, increasing the total MoE parameters consistently enhances performance across all active parameter budgets, and the optimal expert size is determined solely by the number of active parameters. Furthermore, the effectiveness of dropless routing is empirically validated, demonstrating consistent performance gains.
该研究通过调整专家激活数量和概率归一化方法,解决细粒度混合专家模型中减少激活专家数时性能下降的问题。
Task Arithmetic (TA) exhibits limited performance in multi-task fusion for large language models (LLMs), primarily due to insufficient linearity assumptions at the full-model level. Method: We observe that individual model submodules—particularly attention and MLP layers—exhibit significantly higher intrinsic linearity than the global model, and we leverage this as a novel structural prior. Accordingly, we propose Submodule-Level Linear Weighted Merging (SLWM): a closed-form, fine-tuning-free merging method that derives optimal weights per submodule via statistical linearity analysis. Contribution/Results: SLWM extends the TA framework and achieves substantial gains across diverse LLM scales (7B–70B) and multi-task settings (e.g., instruction following + code generation + mathematical reasoning). It consistently outperforms standard TA and other baselines, markedly improving both multi-task generalization and merging stability without additional training or hyperparameter tuning.
研究通过动态专家剪枝方法在细粒度MoE架构中减少冗余专家选择,保留约2/3专家即可保持98.8%性能,提高推理效率。
This study addresses the limitation of static pruning for Mixture-of-Experts (MoE) models in heterogeneous multi-agent systems, where fixed architectures struggle to accommodate dynamic task demands. To this end, we propose a dynamic expert pruning method that employs a lightweight predictor to generate bespoke masks in real time based on system prompts, enabling on-demand expert activation. Notably, this work introduces the first mechanism capable of producing dynamic masks via a single forward pass without requiring offline calibration. Experimental results demonstrate that our approach outperforms static baselines in accuracy across varying model scales and unseen workflows. By significantly reducing the number of retained experts while effectively controlling performance degradation, the proposed method substantially enhances inference serving sparsity and deployment efficiency.
This study addresses the limitation of relying solely on expert utilization rates to assess removal damage during expert pruning in Mixture-of-Experts (MoE) models, emphasizing the necessity of preserving functional substitutability to maintain output distributions. To this end, it proposes a training-free expert pruning framework that introduces a consensus residual-based scoring mechanism for evaluating functional substitutability. By integrating an exact single-deletion identity with calibration token aggregation, the method achieves precise pruning without requiring gradients or recovery training. Extensive experiments across multiple large-scale models and varying pruning ratios demonstrate that the proposed approach attains the highest macro-average score over nine evaluation tasks. It significantly outperforms the REAP baseline while effectively reducing reverse KL divergence, highlighting its efficacy in maintaining model performance under aggressive pruning conditions.
This work addresses the inefficiency in fine-tuning Mixture-of-Experts (MoE) models, which often stems from expert redundancy and uniform parameter allocation, while existing parameter-efficient fine-tuning (PEFT) methods fail to account for routing dynamics. The authors propose EPnG, a novel framework that dynamically couples expert importance with LoRA capacity: it evaluates expert contributions via router gating probabilities, prunes low-importance experts, and orthogonally expands the LoRA rank of high-importance experts. This approach aligns routing with parameter allocation under a fixed budget. Evaluated on OLMoE and Qwen1.5-MoE, EPnG updates only 0.55%–0.72% of parameters—140 to 180 times fewer than full fine-tuning—while matching its performance and significantly outperforming standard LoRA.
This work addresses the high deployment cost of Mixture-of-Experts (MoE) models caused by their massive expert parameters, a challenge inadequately resolved by existing compression methods that struggle to balance accuracy and scalability. The authors propose an efficient compression approach that preserves the original router and leverages functional co-activation patterns among experts to cluster them. Within each cluster, one full-precision dominant expert is retained, while others are represented as low-rank corrections. Furthermore, they introduce BTExperts, a tree-structured organization enabling computation sharing during inference. Evaluated on Qwen3-30B-A3B and Gemma-4-26B-A4B, the method achieves approximately 50% expert compression while outperforming baseline models in downstream accuracy and perplexity across most tasks, with performance gains increasing as the number of experts scales.