dynamic expert expansion

Designs and implements modules and training procedures that dynamically add, grow, prune, and initialize expert submodules within a model to adjust representational capacity without retraining the entire system. This work includes algorithms for deciding when to expand or prune experts, methods for rank growth with orthogonal initialization, and strategies to scale the expert set while maintaining a fixed parameter budget.

dynamicexpertexpansion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses catastrophic forgetting in fine-tuning pretrained models, where newly acquired knowledge overwrites previously learned information. To mitigate this issue, the authors propose a function-preserving model expansion approach that mathematically duplicates and scales parameters of selected Transformer submodules during initialization. This technique enables stable training and faithful retention of original model capabilities without altering the initial functionality. By circumventing the traditional trade-off between plasticity and stability, the method achieves performance comparable to full fine-tuning while expanding only a minimal number of layers. Consequently, it fully preserves the model’s original knowledge and substantially reduces computational overhead.

catastrophic forgettingfine-tuningplasticity-stability trade-off

This work systematically dissects the multidimensional design space of Mixture-of-Experts (MoE) architectures in large language models, moving beyond conventional generational narratives. It introduces a five-dimensional analytical framework encompassing expert granularity, topology, routing flexibility, load-balancing scope, and execution structure, and constructs a dependency graph to elucidate the coupling mechanisms across four control planes: expert topology, routing, load balancing, and expert parallelism. The framework’s validity is empirically demonstrated through iso-budget pretraining experiments integrating algorithmic innovations—such as Top-k routing, shared and fine-grained experts, and dynamic expert composition—with system-level optimizations including token dispatch, device placement, and all-to-all communication. The study further distills key open challenges for the future development of MoE systems.

expert parallelismexpert topologyload balancing

This work systematically investigates the interplay among key design dimensions in Mixture-of-Experts (MoE) architectures—such as the number of experts, expert granularity, heterogeneity, shared experts, and load balancing—through over 2,000 large-scale pretraining experiments. The study reveals that the number of experts and their granularity are the dominant factors governing model performance, while other design choices exert comparatively limited influence. Notably, increasing the total MoE parameters consistently enhances performance across all active parameter budgets, and the optimal expert size is determined solely by the number of active parameters. Furthermore, the effectiveness of dropless routing is empirically validated, demonstrating consistent performance gains.

expert countexpert granularityload balancing

Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMs

Apr 15, 2025
RD
Rui Dai
🏛️ University of Science and Technology of China | Independent Researcher | Hong Kong Baptist University

Task Arithmetic (TA) exhibits limited performance in multi-task fusion for large language models (LLMs), primarily due to insufficient linearity assumptions at the full-model level. Method: We observe that individual model submodules—particularly attention and MLP layers—exhibit significantly higher intrinsic linearity than the global model, and we leverage this as a novel structural prior. Accordingly, we propose Submodule-Level Linear Weighted Merging (SLWM): a closed-form, fine-tuning-free merging method that derives optimal weights per submodule via statistical linearity analysis. Contribution/Results: SLWM extends the TA framework and achieves substantial gains across diverse LLM scales (7B–70B) and multi-task settings (e.g., instruction following + code generation + mathematical reasoning). It consistently outperforms standard TA and other baselines, markedly improving both multi-task generalization and merging stability without additional training or hyperparameter tuning.

Enhancing task arithmetic performance via submodule linearity in LLMsImproving multi-task capabilities through independent submodule mergingProposing a model merging strategy using submodules' linear properties

Latest Papers

What's happening recently
View more

This study addresses the limitation of static pruning for Mixture-of-Experts (MoE) models in heterogeneous multi-agent systems, where fixed architectures struggle to accommodate dynamic task demands. To this end, we propose a dynamic expert pruning method that employs a lightweight predictor to generate bespoke masks in real time based on system prompts, enabling on-demand expert activation. Notably, this work introduces the first mechanism capable of producing dynamic masks via a single forward pass without requiring offline calibration. Experimental results demonstrate that our approach outperforms static baselines in accuracy across varying model scales and unseen workflows. By significantly reducing the number of retained experts while effectively controlling performance degradation, the proposed method substantially enhances inference serving sparsity and deployment efficiency.

Dynamic PruningExpert PruningMemory Efficiency

This study addresses the limitation of relying solely on expert utilization rates to assess removal damage during expert pruning in Mixture-of-Experts (MoE) models, emphasizing the necessity of preserving functional substitutability to maintain output distributions. To this end, it proposes a training-free expert pruning framework that introduces a consensus residual-based scoring mechanism for evaluating functional substitutability. By integrating an exact single-deletion identity with calibration token aggregation, the method achieves precise pruning without requiring gradients or recovery training. Extensive experiments across multiple large-scale models and varying pruning ratios demonstrate that the proposed approach attains the highest macro-average score over nine evaluation tasks. It significantly outperforms the REAP baseline while effectively reducing reverse KL divergence, highlighting its efficacy in maintaining model performance under aggressive pruning conditions.

Expert PruningFunctional ReplaceabilityLarge Language Models

This work addresses the inefficiency in fine-tuning Mixture-of-Experts (MoE) models, which often stems from expert redundancy and uniform parameter allocation, while existing parameter-efficient fine-tuning (PEFT) methods fail to account for routing dynamics. The authors propose EPnG, a novel framework that dynamically couples expert importance with LoRA capacity: it evaluates expert contributions via router gating probabilities, prunes low-importance experts, and orthogonally expands the LoRA rank of high-importance experts. This approach aligns routing with parameter allocation under a fixed budget. Evaluated on OLMoE and Qwen1.5-MoE, EPnG updates only 0.55%–0.72% of parameters—140 to 180 times fewer than full fine-tuning—while matching its performance and significantly outperforming standard LoRA.

expert redundancyMixture-of-Expertsparameter allocation

This work addresses the high deployment cost of Mixture-of-Experts (MoE) models caused by their massive expert parameters, a challenge inadequately resolved by existing compression methods that struggle to balance accuracy and scalability. The authors propose an efficient compression approach that preserves the original router and leverages functional co-activation patterns among experts to cluster them. Within each cluster, one full-precision dominant expert is retained, while others are represented as low-rank corrections. Furthermore, they introduce BTExperts, a tree-structured organization enabling computation sharing during inference. Evaluated on Qwen3-30B-A3B and Gemma-4-26B-A4B, the method achieves approximately 50% expert compression while outperforming baseline models in downstream accuracy and perplexity across most tasks, with performance gains increasing as the number of experts scales.

expert compressionlow-rank decompositionMixture-of-Experts

Hot Scholars

NV

Nalini Venkatasubramanian

Professor of Computer Science, University of California, Irvine
distributed systemsmiddlewareInternet-of-Thingscyberphysical systems
KR

K. R. Jayaram

Research Scientist, IBM Research
Distributed SystemsProgramming Languages
HY

Hongzhi Yin

Professor and ARC Future Fellow, University of Queensland
Recommender SystemGraph LearningSpatial-temporal PredictionEdge Intelligence