Score
Design and implement procedures to adjust and realign Mixture-of-Experts (MoE) router/gating parameters so that expert assignment probabilities remain calibrated after operations like merging or transferring routers. This includes deriving and applying Hessian/second-order-curvature-aware (HARC) calibration algorithms, closed-form or efficient realignment solutions, and practical steps to mitigate routing breakdown without full retraining.
Efficiently reusing multiple domain- or task-specific fine-tuned expert models while achieving high performance and strong generalization remains challenging. Method: We propose MoErging—a unified methodology for model merging, Mixture of Experts (MoE), and multi-task learning—featuring input-aware dynamic routing, parameter-space fusion, learnable router design, and collaborative multi-expert inference. Contribution/Results: We introduce the first taxonomy of MoErging methods; develop an open-source toolchain and standardized evaluation benchmark; and construct the first multidimensional MoErging knowledge graph. Our analysis rigorously characterizes applicability boundaries and performance trade-offs across paradigms, establishing a theoretical framework and practical guidelines for collaborative model reuse.
This work addresses the sensitivity of Mixture-of-Experts (MoE) models to parameter perturbations during model merging, which often leads to routing collapse and severe performance degradation. The study is the first to identify this issue and introduces Hessian-aware Router Calibration (HARC), a training-free framework that leverages second-order curvature information from the Hessian matrix to analytically realign the router post-merging, thereby restoring its routing capability. By integrating matrix-free conjugate gradient methods with Top-k routing analysis, HARC significantly enhances the performance of various MoE merging baselines on mathematical reasoning and code generation tasks, effectively mitigating routing failure without additional training.
This work addresses the lack of effective design principles for routers in existing Mixture-of-Experts (MoE) models, which struggle to accurately capture the affinity between tokens and experts. The authors propose, for the first time, using the dominant singular directions of expert matrices as the target for router design and introduce a novel “power iteration followed by shrinkage” paradigm. During pretraining, they employ manifold optimization to dynamically align the router’s row vectors with these dominant singular directions. This approach achieves a favorable balance among alignment accuracy, computational efficiency, and training stability. Experiments on MoE models ranging from 1B to 11B parameters demonstrate substantial performance improvements, validating the effectiveness of the proposed router redesign strategy.
研究解决了MoE模型后训练中保持基础路由结构的问题,通过引入Router Prior Bias方法,使下游任务性能提升并保留更多领域外能力。
To address suboptimal expert routing in Mixture-of-Experts (MoE) models during deployment under distribution shift, this paper proposes a training-free, data-free online adaptive routing framework. The method dynamically re-routes experts during autoregressive generation based on the already-produced sequence, optimizing router logits via periodic self-supervised updates—only lightweight learnable vectors are adjusted to mitigate overfitting. This constitutes the first plug-and-play, reference-free online routing adaptation for MoE models. Evaluated on OLMoE, it improves HumanEval scores by 5.5%; when integrated with DeepSeek-V2-Lite and self-consistency decoding, it yields an average 6% performance gain, significantly enhancing generation robustness and consistency under complex reasoning and distribution-shift scenarios.
This work addresses the slow statistical convergence of cosine routers in Mixture-of-Experts (MoE) models. We first establish—via theoretical analysis—that parameter estimation under standard cosine routing exhibits only logarithmic convergence due to weak identifiability of routing directions and non-identifiability of input ℓ² norms. To resolve this, we propose a perturbed cosine router that injects controlled noise into the input ℓ² norm, breaking symmetry-induced degeneracy and restoring strong identifiability. We prove that this perturbation elevates both expert weight and function estimation convergence to polynomial rates (O(1/n^α), α > 0), thereby filling a critical gap in the statistical analysis of MoE routing mechanisms. Our approach integrates least-squares estimation, PDE-based modeling, and principled perturbation design. Empirical validation on synthetic data as well as image and language tasks demonstrates substantial mitigation of representation collapse and consistent improvements in generalization performance.
This work addresses the challenge of efficiently compressing Mixture-of-Experts (MoE) models, which typically require loading all expert parameters. The authors propose a one-shot expert pruning method based on lightweight fine-tuning—such as router-specific LoRA or IA³—that induces changes in router weights. By measuring the ℓ² norm of these weight changes to assess expert sensitivity, they rank and prune the least sensitive experts. This study is the first to demonstrate that router sensitivity serves as an effective pruning signal, enabling near-linear accuracy degradation rather than catastrophic collapse under high compression ratios, with notable transferability across models. On Mixtral-8×7B, pruning 50% of experts yields a 28.76% score on MMLU-Pro, alongside 49% memory reduction and 37% lower latency; on Qwen1.5-MoE, it maintains a 49.7% average accuracy on mathematical tasks, substantially outperforming random or magnitude-based pruning.
This work addresses a key limitation in sparse Mixture-of-Experts (MoE) models, where the routing mechanism jointly handles expert selection and output weighting, potentially constraining performance. The study provides the first systematic validation that these two functions should be decoupled and introduces Fixed Dispatch with Adaptive Aggregation (FDAA): a lightweight, learnable aggregation head is added atop a frozen backbone and fixed expert assignments, enabling end-to-end optimization of aggregation weights via the language modeling objective. Evaluated on pretrained MoE models such as OLMoE and DeepSeek-V2-Lite, FDAA achieves a 0.1523 reduction in cross-entropy on WikiText-103 and demonstrates consistent improvements across diverse benchmarks including C4 and PTB, confirming both the efficacy and generality of the proposed decoupling strategy.
This work addresses a critical limitation in existing dynamic routing methods, which conflate irreducible ambiguity with recoverable risk that can be mitigated by ensembling more experts. The authors formalize routing as an information-value allocation problem, introducing a novel mechanism that generates simultaneous upper-bound risk certificates via counterfactual risk estimation. By greedily allocating computational budget based on marginal risk reduction per unit cost, the method decides whether to answer or abstain. This approach is the first to explicitly disentangle the two types of uncertainty and, when integrated with a LoRA-based mixture-of-experts architecture, provides theoretical guarantees on risk certification and optimal resource allocation. Experiments demonstrate significant accuracy gains under identical computational budgets, effective high-coverage risk control, and superior performance over current MoE-LoRA baselines under distribution shifts, tail latency, and risk-coverage trade-offs.
This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.
This work addresses the limitation of fixed-expert MoE-LoRA architectures, which inefficiently allocate computation by employing a constant number of experts regardless of input token difficulty—wasting resources on easy tokens while under-provisioning for challenging ones. To overcome this, the authors propose CARE, a novel dynamic routing method that leverages both the confidence of the router’s output distribution and inter-expert disagreement as signals to guide expert selection. Experts are activated via nucleus sampling until their cumulative weights reach an adaptive threshold, while a budget thermostat regulates the average number of active experts. Requiring no additional parameters and only a single forward pass, CARE achieves comparable or superior performance to top-k=4 MoE-LoRA on LLaMA-3.1-8B and Qwen2.5-7B using fewer experts, while also significantly enhancing out-of-distribution detection capability.