calibrate moe routers

Design and implement procedures to adjust and realign Mixture-of-Experts (MoE) router/gating parameters so that expert assignment probabilities remain calibrated after operations like merging or transferring routers. This includes deriving and applying Hessian/second-order-curvature-aware (HARC) calibration algorithms, closed-form or efficient realignment solutions, and practical steps to mitigate routing breakdown without full retraining.

calibratemoerouters

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.59
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the sensitivity of Mixture-of-Experts (MoE) models to parameter perturbations during model merging, which often leads to routing collapse and severe performance degradation. The study is the first to identify this issue and introduces Hessian-aware Router Calibration (HARC), a training-free framework that leverages second-order curvature information from the Hessian matrix to analytically realign the router post-merging, thereby restoring its routing capability. By integrating matrix-free conjugate gradient methods with Top-k routing analysis, HARC significantly enhances the performance of various MoE merging baselines on mathematical reasoning and code generation tasks, effectively mitigating routing failure without additional training.

Mixture-of-ExpertsModel MergingMoE

This work addresses the lack of effective design principles for routers in existing Mixture-of-Experts (MoE) models, which struggle to accurately capture the affinity between tokens and experts. The authors propose, for the first time, using the dominant singular directions of expert matrices as the target for router design and introduce a novel “power iteration followed by shrinkage” paradigm. During pretraining, they employ manifold optimization to dynamically align the router’s row vectors with these dominant singular directions. This approach achieves a favorable balance among alignment accuracy, computational efficiency, and training stability. Experiments on MoE models ranging from 1B to 11B parameters demonstrate substantial performance improvements, validating the effectiveness of the proposed router redesign strategy.

expert representationMixture-of-Expertsrouter design

Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models

Oct 16, 2025
GS
Guinan Su
🏛️ Max Planck Institute for Intelligent Systems | University of Tübingen | Sun Yat-sen University | University of Surrey | ELLIS Institute Tübingen | Töbingen AI Center

To address suboptimal expert routing in Mixture-of-Experts (MoE) models during deployment under distribution shift, this paper proposes a training-free, data-free online adaptive routing framework. The method dynamically re-routes experts during autoregressive generation based on the already-produced sequence, optimizing router logits via periodic self-supervised updates—only lightweight learnable vectors are adjusted to mitigate overfitting. This constitutes the first plug-and-play, reference-free online routing adaptation for MoE models. Evaluated on OLMoE, it improves HumanEval scores by 5.5%; when integrated with DeepSeek-V2-Lite and self-consistency decoding, it yields an average 6% performance gain, significantly enhancing generation robustness and consistency under complex reasoning and distribution-shift scenarios.

Addressing suboptimal expert selection due to distribution shiftsEnabling data-free online adaptation without external supervisionOptimizing MoE routing decisions dynamically during text generation

Statistical Advantages of Perturbing Cosine Router in Mixture of Experts

May 23, 2024
HN
Huy Nguyen
🏛️ The University of Texas at Austin | VinAI Research

This work addresses the slow statistical convergence of cosine routers in Mixture-of-Experts (MoE) models. We first establish—via theoretical analysis—that parameter estimation under standard cosine routing exhibits only logarithmic convergence due to weak identifiability of routing directions and non-identifiability of input ℓ² norms. To resolve this, we propose a perturbed cosine router that injects controlled noise into the input ℓ² norm, breaking symmetry-induced degeneracy and restoring strong identifiability. We prove that this perturbation elevates both expert weight and function estimation convergence to polynomial rates (O(1/n^α), α > 0), thereby filling a critical gap in the statistical analysis of MoE routing mechanisms. Our approach integrates least-squares estimation, PDE-based modeling, and principled perturbation design. Empirical validation on synthetic data as well as image and language tasks demonstrates substantial mitigation of representation collapse and consistent improvements in generalization performance.

Analyzes slow convergence rates in cosine router MoE.Proposes perturbed cosine router to improve estimation rates.Validates improved polynomial convergence rates empirically.

Latest Papers

What's happening recently
View more

This work addresses the challenge of efficiently compressing Mixture-of-Experts (MoE) models, which typically require loading all expert parameters. The authors propose a one-shot expert pruning method based on lightweight fine-tuning—such as router-specific LoRA or IA³—that induces changes in router weights. By measuring the ℓ² norm of these weight changes to assess expert sensitivity, they rank and prune the least sensitive experts. This study is the first to demonstrate that router sensitivity serves as an effective pruning signal, enabling near-linear accuracy degradation rather than catastrophic collapse under high compression ratios, with notable transferability across models. On Mixtral-8×7B, pruning 50% of experts yields a 28.76% score on MMLU-Pro, alongside 49% memory reduction and 37% lower latency; on Qwen1.5-MoE, it maintains a 49.7% average accuracy on mathematical tasks, substantially outperforming random or magnitude-based pruning.

expert pruninglightweight fine-tuningMixture-of-Experts

This work addresses a key limitation in sparse Mixture-of-Experts (MoE) models, where the routing mechanism jointly handles expert selection and output weighting, potentially constraining performance. The study provides the first systematic validation that these two functions should be decoupled and introduces Fixed Dispatch with Adaptive Aggregation (FDAA): a lightweight, learnable aggregation head is added atop a frozen backbone and fixed expert assignments, enabling end-to-end optimization of aggregation weights via the language modeling objective. Evaluated on pretrained MoE models such as OLMoE and DeepSeek-V2-Lite, FDAA achieves a 0.1523 reduction in cross-entropy on WikiText-103 and demonstrates consistent improvements across diverse benchmarks including C4 and PTB, confirming both the efficacy and generality of the proposed decoupling strategy.

expert aggregationexpert dispatchMixture-of-Experts

This work addresses a critical limitation in existing dynamic routing methods, which conflate irreducible ambiguity with recoverable risk that can be mitigated by ensembling more experts. The authors formalize routing as an information-value allocation problem, introducing a novel mechanism that generates simultaneous upper-bound risk certificates via counterfactual risk estimation. By greedily allocating computational budget based on marginal risk reduction per unit cost, the method decides whether to answer or abstain. This approach is the first to explicitly disentangle the two types of uncertainty and, when integrated with a LoRA-based mixture-of-experts architecture, provides theoretical guarantees on risk certification and optimal resource allocation. Experiments demonstrate significant accuracy gains under identical computational budgets, effective high-coverage risk control, and superior performance over current MoE-LoRA baselines under distribution shifts, tail latency, and risk-coverage trade-offs.

mixture of expertsrisk certificationrouting

This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.

dynamic routingexpert capacityMixture-of-Experts

This work addresses the limitation of fixed-expert MoE-LoRA architectures, which inefficiently allocate computation by employing a constant number of experts regardless of input token difficulty—wasting resources on easy tokens while under-provisioning for challenging ones. To overcome this, the authors propose CARE, a novel dynamic routing method that leverages both the confidence of the router’s output distribution and inter-expert disagreement as signals to guide expert selection. Experts are activated via nucleus sampling until their cumulative weights reach an adaptive threshold, while a budget thermostat regulates the average number of active experts. Requiring no additional parameters and only a single forward pass, CARE achieves comparable or superior performance to top-k=4 MoE-LoRA on LLaMA-3.1-8B and Qwen2.5-7B using fewer experts, while also significantly enhancing out-of-distribution detection capability.

Confidence-Adaptive RoutingExpert AllocationLow-Rank Adaptation

Hot Scholars

JW

Justus Westerhoff

Ph.D. Student, Berliner Hochschule für Technik (BHT)
Bio-Inspired AIContinual Learning
JZ

Jianhua Zhang

Beijing University of Posts and Telecommunications, CHINA
Signal ProcessingWireless CommunicationRadio channel Measurement and ModellingChannel Simulation
HP

Hitesh Poddar

Sharp Laboratories of America, Inc., NYU WIRELESS
6Gwirelesspropagation