Score
Design, implement, or analyze routing modules that perform discrete, token-level top-1 expert selection (assigning exactly one expert per token), including the mechanics of hard selection and token-level expert assignment. This includes implementing training techniques such as straight-through estimators to enable gradient flow through the discrete choice, building lightweight shared routers, and enforcing constraints that keep each expert’s contribution at a unit scale.
Efficiently reusing multiple domain- or task-specific fine-tuned expert models while achieving high performance and strong generalization remains challenging. Method: We propose MoErging—a unified methodology for model merging, Mixture of Experts (MoE), and multi-task learning—featuring input-aware dynamic routing, parameter-space fusion, learnable router design, and collaborative multi-expert inference. Contribution/Results: We introduce the first taxonomy of MoErging methods; develop an open-source toolchain and standardized evaluation benchmark; and construct the first multidimensional MoErging knowledge graph. Our analysis rigorously characterizes applicability boundaries and performance trade-offs across paradigms, establishing a theoretical framework and practical guidelines for collaborative model reuse.
This work addresses a key limitation in sparse Mixture-of-Experts (MoE) models, where the routing mechanism jointly handles expert selection and output weighting, potentially constraining performance. The study provides the first systematic validation that these two functions should be decoupled and introduces Fixed Dispatch with Adaptive Aggregation (FDAA): a lightweight, learnable aggregation head is added atop a frozen backbone and fixed expert assignments, enabling end-to-end optimization of aggregation weights via the language modeling objective. Evaluated on pretrained MoE models such as OLMoE and DeepSeek-V2-Lite, FDAA achieves a 0.1523 reduction in cross-entropy on WikiText-103 and demonstrates consistent improvements across diverse benchmarks including C4 and PTB, confirming both the efficacy and generality of the proposed decoupling strategy.
This work proposes a novel mixture-of-experts (MoE) architecture that eliminates the need for explicit routing mechanisms commonly found in traditional MoE models. By embedding activation logic directly within each expert and enabling end-to-end continuous gradient flow, experts autonomously determine their own activation without reliance on external routers, Softmax operations, Top-K selection, or hard-coded load-balancing heuristics. The approach introduces a unified, adaptive load-balancing framework that jointly optimizes resource allocation across both experts and tokens, supporting configurable dual-objective balancing. Experimental results demonstrate that the proposed model consistently outperforms existing baselines across multiple benchmarks, exhibiting superior scalability and robustness while removing rigid inductive biases imposed by centralized routing.
This work addresses the suboptimality of standard top-k routing in Mixture-of-Experts language models for complex reasoning tasks, where direct evaluation of routing efficacy has been lacking. By freezing model parameters and comparing standard routing against counterfactual alternatives with equivalent computational cost—using next-token prediction probabilities along ground-truth reasoning trajectories as a utility metric—the study reveals that routers perform well on high-confidence tokens but fail at fragile reasoning steps. This limitation stems from training objectives that optimize only the executed path and rely on statistical load balancing. To mitigate this, the authors propose fine-tuning only the final-layer router, which significantly improves pass@K performance on AIME 2024+2025 and HMMT 2025 benchmarks in Qwen3-30B-A3B and GPT-OSS-20B.
This work uncovers the dynamic trade-off between load balancing and model quality in Mixture-of-Experts (MoE) training. By modeling token routing as a congestion game and introducing an effective congestion coefficient γ_eff to monitor the entire training trajectory, the study reveals—for the first time—a non-monotonic three-phase evolution of load balancing: surge, stabilization, and relaxation. This pattern elucidates the intrinsic mechanism of prioritizing balancing early in training and model quality later. The approach integrates temperature-scaled softmax, multi-type congestion decomposition, token clustering, and range diagnostics (K/M, ε_l). Evaluated on OLMoE-1B–7B and OpenMoE-8B, it achieves a 30% average improvement in load prediction, model quality estimation consistency with Pearson correlation r ≥ 0.89, and L1 error approaching the theoretical lower bound.
This study addresses the insufficient expert specialization in sparse Mixture-of-Experts models caused by conflicts between load balancing and gradient alignment. To mitigate this, we propose Gradient-Aligned Routing (GAR), which reformulates routing as a gradient partitioning problem. By decoupling load balancing from gradient-based routing organization, GAR rewards samples with consistent gradients within groups to optimize multi-task classification performance. We validate our method across five multi-task text classification datasets using RoBERTa, DeBERTa, and Qwen3 backbones integrated with low-rank adapter experts. Experimental results demonstrate that GAR improves accuracy by approximately 1.1 percentage points over task-loss-only routing while significantly enhancing gradient purity. These findings highlight the critical value of leveraging gradient information for strengthening expert specialization.
This work addresses the inefficiency of conventional Mixture-of-Experts (MoE) models, which employ a fixed top-k expert selection strategy that fails to dynamically allocate computational resources according to individual token demands. To overcome this limitation, the authors propose a training-free, plug-in method for inference that introduces, for the first time in MoE architectures, an elbow-point detection mechanism. By analyzing the probability distribution output by the router, this approach adaptively determines the number of experts to activate per token. Integrating principles from ranking and load balancing theory, the method achieves dynamic resource allocation while preserving balanced expert utilization. Experimental results demonstrate that the proposed technique reduces average inference latency by 5.3% across mainstream MoE models without compromising accuracy on six benchmark evaluations.
This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.
This study addresses the challenge of optimizing non-additive rewards in online routing over expert subsets with bilateral constraints. We propose the Multi-Subset Routing (MSR) framework and an OMD-Approachability algorithm that integrates Online Mirror Descent with Blackwell’s approachability theory. This method overcomes the additive reward assumption inherent in traditional combinatorial bandits, achieving both constraint satisfaction and coverage maximization under winner-only feedback. Theoretical analysis demonstrates that both regret and constraint violation are bounded by O(1/√T). Furthermore, empirical evaluations on real-world crowdsourcing datasets validate the effectiveness of our approach. Collectively, this work establishes a novel paradigm for non-additive combinatorial online learning, extending the applicability of online optimization to complex routing scenarios where reward structures are inherently non-linear and constrained.
This study addresses the limitations of flat routing in conventional Mixture-of-Experts (MoE) models, which often suffer from expert load imbalance and a lack of topological structure. To overcome these issues, this work proposes a hierarchical binary decision tree routing architecture that dynamically estimates branch probabilities via exponential moving averages to balance traffic across subtrees. This mechanism achieves load balancing without auxiliary losses while theoretically preventing routing collapse. Experimental results demonstrate that the proposed method preserves task accuracy across multiple benchmark datasets while significantly reducing cross-device communication overhead in distributed settings, thereby enabling efficient and balanced expert utilization.
This study addresses the substantial memory overhead of deploying Mixture-of-Experts (MoE) models and the limitations of existing pruning methods, including poor alignment, high computational cost, and neglect of routing redundancy. We propose an efficient MoE pruning framework that achieves structured pruning by introducing learnable router biases and diversity regularization to precisely identify critical experts. Furthermore, an affine transformation-based expert approximation mechanism is designed to effectively compensate for accuracy degradation caused by pruning. Experimental results demonstrate that the proposed method successfully removes 25%–50% of experts across multiple large language models while consistently outperforming state-of-the-art algorithms on nine zero-shot benchmarks, achieving a favorable balance between model compression and reasoning capability.