Score
Designs and implements methods that combine multiple expert parameter sets into compact, mergeable representations or runtime selectors by representing experts as lightweight ‘‘task’’ vectors and using techniques such as low‑rank residual decompositions, subspace alignment, learned or mask‑based parameter selection, and prototype- or manifold‑based routing. Builds merging operators and runtime policies that avoid storing full expert parameters, enable parameter‑efficient deployment and dynamic per‑input expert selection, and analyzes tradeoffs in memory, compute, and task‑retention to preserve task knowledge adaptively.
This work systematically dissects the multidimensional design space of Mixture-of-Experts (MoE) architectures in large language models, moving beyond conventional generational narratives. It introduces a five-dimensional analytical framework encompassing expert granularity, topology, routing flexibility, load-balancing scope, and execution structure, and constructs a dependency graph to elucidate the coupling mechanisms across four control planes: expert topology, routing, load balancing, and expert parallelism. The framework’s validity is empirically demonstrated through iso-budget pretraining experiments integrating algorithmic innovations—such as Top-k routing, shared and fine-grained experts, and dynamic expert composition—with system-level optimizations including token dispatch, device placement, and all-to-all communication. The study further distills key open challenges for the future development of MoE systems.
This work addresses the co-occurring challenges of backbone parameter saturation and expert redundancy in continual model merging, as well as bottlenecks induced by data-driven routing. To this end, the authors propose MADE-IT, a novel approach featuring a manifold-aware dynamic expert evolution mechanism that autonomously adds or removes experts based on projection subspace affinity and a distribution-aware adaptive threshold. Additionally, MADE-IT introduces a data-free, training-free implicit routing strategy that guides expert activation through feature–subspace alignment. Experimental results demonstrate that MADE-IT significantly outperforms baseline methods on long-sequence and out-of-order tasks, achieving higher accuracy and robustness while substantially reducing expert redundancy—particularly within general-purpose modules and shallow network layers.
This work addresses the performance degradation of task-specific experts in multitask model merging caused by parameter interference, as well as the high inference cost and storage overhead of existing dynamic methods that rely on redundant expert copies. The authors propose ReTeX, a framework that models parameter interference as an affine transformation of expert parameters and approximates it with a learnable additive offset, enabling a single merged model to recover near-original expert performance. Innovatively, ReTeX introduces a router-free task identifier that leverages singular value decomposition (SVD) subspace projection residuals to match task identities, achieving the first subspace-based task recognition without additional storage. Experiments demonstrate that ReTeX recovers over 95% of standalone expert performance across vision and NLP tasks and exhibits strong generalization and adaptive knowledge interpolation capabilities on unseen tasks.
This work addresses the performance degradation commonly observed in merged multi-task models due to parameter interference, which often results in inferior performance compared to single-task experts. Existing dynamic routing approaches typically require additional training or prior knowledge of task identities, limiting their practicality. To overcome these limitations, the authors propose a training-free, task-ID-agnostic dynamic routing mechanism that leverages a few task-specific support samples to construct low-rank task manifolds via singular value decomposition (SVD). Routing decisions are made by evaluating the projection residuals of test samples onto these manifolds. The method seamlessly integrates with lightweight subspace- or mask-based merging strategies and demonstrates consistent performance gains across multiple computer vision and natural language processing benchmarks, effectively narrowing the gap with single-task expert models even when task identities are unknown at inference time.
This work investigates the opaque expert specialization mechanism in Mixture-of-Experts (MoE) models, which limits inference and memory efficiency. By analyzing domain-specific routing patterns and employing an early-decoding framework, the study systematically examines how individual experts contribute to model outputs. Through comprehensive analyses—including routing distribution statistics, cosine similarity of hidden states, comparisons between single-expert and ensemble outputs, and perplexity evaluation—the authors find that a small subset of experts handles over 50% of all requests. Remarkably, outputs from a single dominant expert exhibit high consistency with the full model (cosine similarity up to 0.95), with only a 5% increase in perplexity. These findings suggest that precise expert pruning can substantially enhance inference efficiency without compromising performance, offering a promising avenue for efficient MoE deployment and knowledge localization.
This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.
This work addresses the high deployment cost of Mixture-of-Experts (MoE) models caused by their massive expert parameters, a challenge inadequately resolved by existing compression methods that struggle to balance accuracy and scalability. The authors propose an efficient compression approach that preserves the original router and leverages functional co-activation patterns among experts to cluster them. Within each cluster, one full-precision dominant expert is retained, while others are represented as low-rank corrections. Furthermore, they introduce BTExperts, a tree-structured organization enabling computation sharing during inference. Evaluated on Qwen3-30B-A3B and Gemma-4-26B-A4B, the method achieves approximately 50% expert compression while outperforming baseline models in downstream accuracy and perplexity across most tasks, with performance gains increasing as the number of experts scales.
This work proposes the DS-MoE framework to address the memory bottlenecks in deploying Mixture-of-Experts (MoE) models and the limitation of conventional Top-k routing in neglecting expert dependencies. The framework reveals the functional duality of expert combinations and formulates the selection objective as a difference-of-submodular function, thereby decoupling redundancy from synergy effects. Furthermore, based on second-order Taylor expansion analysis, it designs a majorization-minimization algorithm with monotonicity guarantees to achieve dependency-aware extraction of compact expert subsets. Experimental results demonstrate that the proposed method effectively preserves critical expert combinations and significantly outperforms existing baselines.
This work addresses the inefficiency of conventional Mixture-of-Experts (MoE) models, which employ a fixed top-k expert selection strategy that fails to dynamically allocate computational resources according to individual token demands. To overcome this limitation, the authors propose a training-free, plug-in method for inference that introduces, for the first time in MoE architectures, an elbow-point detection mechanism. By analyzing the probability distribution output by the router, this approach adaptively determines the number of experts to activate per token. Integrating principles from ranking and load balancing theory, the method achieves dynamic resource allocation while preserving balanced expert utilization. Experimental results demonstrate that the proposed technique reduces average inference latency by 5.3% across mainstream MoE models without compromising accuracy on six benchmark evaluations.
本文通过局部聚合视角分析混合专家模型,探讨路由、稀疏激活和共享专家等设计选择的统计作用,分离出逼近误差、专家学习误差和路由器估计误差。