hard top-1 routing

Design, implement, or analyze routing modules that perform discrete, token-level top-1 expert selection (assigning exactly one expert per token), including the mechanics of hard selection and token-level expert assignment. This includes implementing training techniques such as straight-through estimators to enable gradient flow through the discrete choice, building lightweight shared routers, and enforcing constraints that keep each expert’s contribution at a unit scale.

hardtop-1routing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.79
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a key limitation in sparse Mixture-of-Experts (MoE) models, where the routing mechanism jointly handles expert selection and output weighting, potentially constraining performance. The study provides the first systematic validation that these two functions should be decoupled and introduces Fixed Dispatch with Adaptive Aggregation (FDAA): a lightweight, learnable aggregation head is added atop a frozen backbone and fixed expert assignments, enabling end-to-end optimization of aggregation weights via the language modeling objective. Evaluated on pretrained MoE models such as OLMoE and DeepSeek-V2-Lite, FDAA achieves a 0.1523 reduction in cross-entropy on WikiText-103 and demonstrates consistent improvements across diverse benchmarks including C4 and PTB, confirming both the efficacy and generality of the proposed decoupling strategy.

expert aggregationexpert dispatchMixture-of-Experts

This work proposes a novel mixture-of-experts (MoE) architecture that eliminates the need for explicit routing mechanisms commonly found in traditional MoE models. By embedding activation logic directly within each expert and enabling end-to-end continuous gradient flow, experts autonomously determine their own activation without reliance on external routers, Softmax operations, Top-K selection, or hard-coded load-balancing heuristics. The approach introduces a unified, adaptive load-balancing framework that jointly optimizes resource allocation across both experts and tokens, supporting configurable dual-objective balancing. Experimental results demonstrate that the proposed model consistently outperforms existing baselines across multiple benchmarks, exhibiting superior scalability and robustness while removing rigid inductive biases imposed by centralized routing.

expert activationinductive biasload balancing

This work addresses the suboptimality of standard top-k routing in Mixture-of-Experts language models for complex reasoning tasks, where direct evaluation of routing efficacy has been lacking. By freezing model parameters and comparing standard routing against counterfactual alternatives with equivalent computational cost—using next-token prediction probabilities along ground-truth reasoning trajectories as a utility metric—the study reveals that routers perform well on high-confidence tokens but fail at fragile reasoning steps. This limitation stems from training objectives that optimize only the executed path and rely on statistical load balancing. To mitigate this, the authors propose fine-tuning only the final-layer router, which significantly improves pass@K performance on AIME 2024+2025 and HMMT 2025 benchmarks in Qwen3-30B-A3B and GPT-OSS-20B.

expert allocationlanguage modelsMixture-of-Experts

This work uncovers the dynamic trade-off between load balancing and model quality in Mixture-of-Experts (MoE) training. By modeling token routing as a congestion game and introducing an effective congestion coefficient γ_eff to monitor the entire training trajectory, the study reveals—for the first time—a non-monotonic three-phase evolution of load balancing: surge, stabilization, and relaxation. This pattern elucidates the intrinsic mechanism of prioritizing balancing early in training and model quality later. The approach integrates temperature-scaled softmax, multi-type congestion decomposition, token clustering, and range diagnostics (K/M, ε_l). Evaluated on OLMoE-1B–7B and OpenMoE-8B, it achieves a 30% average improvement in load prediction, model quality estimation consistency with Pearson correlation r ≥ 0.89, and L1 error approaching the theoretical lower bound.

congestion gameexpert routingload balancing

This study addresses the insufficient expert specialization in sparse Mixture-of-Experts models caused by conflicts between load balancing and gradient alignment. To mitigate this, we propose Gradient-Aligned Routing (GAR), which reformulates routing as a gradient partitioning problem. By decoupling load balancing from gradient-based routing organization, GAR rewards samples with consistent gradients within groups to optimize multi-task classification performance. We validate our method across five multi-task text classification datasets using RoBERTa, DeBERTa, and Qwen3 backbones integrated with low-rank adapter experts. Experimental results demonstrate that GAR improves accuracy by approximately 1.1 percentage points over task-loss-only routing while significantly enhancing gradient purity. These findings highlight the critical value of leveraging gradient information for strengthening expert specialization.

Expert SpecializationGradient-aligned RoutingMulti-task Learning

Latest Papers

What's happening recently
View more

This work addresses the inefficiency of conventional Mixture-of-Experts (MoE) models, which employ a fixed top-k expert selection strategy that fails to dynamically allocate computational resources according to individual token demands. To overcome this limitation, the authors propose a training-free, plug-in method for inference that introduces, for the first time in MoE architectures, an elbow-point detection mechanism. By analyzing the probability distribution output by the router, this approach adaptively determines the number of experts to activate per token. Integrating principles from ranking and load balancing theory, the method achieves dynamic resource allocation while preserving balanced expert utilization. Experimental results demonstrate that the proposed technique reduces average inference latency by 5.3% across mainstream MoE models without compromising accuracy on six benchmark evaluations.

compute allocationdynamic routingexpert selection

This work addresses the disconnect in existing Mixture-of-Experts (MoE) models between shared computation and dynamic routing, which overlooks the interdependence between reusable computation and residual expert requirements. The paper proposes UniF-MoE, a unified framework introducing a novel “shared-first, routed-later” mechanism: it first processes common features through a shared general-purpose module and then dynamically activates residual experts based on a shared-demand score and complementarity. Key innovations include key prototype selection, cumulative routing quality allocation, and Gram regularization to enhance routing sparsity and diversity, revealing a negative correlation between shared coverage and residual demand. Experiments demonstrate that UniF-MoE outperforms both static and dynamic MoE approaches on DomainBed and GLUE benchmarks while significantly reducing activated computation, inference latency, and memory footprint.

dynamic routingexpert capacityMixture-of-Experts

This study addresses the challenge of optimizing non-additive rewards in online routing over expert subsets with bilateral constraints. We propose the Multi-Subset Routing (MSR) framework and an OMD-Approachability algorithm that integrates Online Mirror Descent with Blackwell’s approachability theory. This method overcomes the additive reward assumption inherent in traditional combinatorial bandits, achieving both constraint satisfaction and coverage maximization under winner-only feedback. Theoretical analysis demonstrates that both regret and constraint violation are bounded by O(1/√T). Furthermore, empirical evaluations on real-world crowdsourcing datasets validate the effectiveness of our approach. Collectively, this work establishes a novel paradigm for non-additive combinatorial online learning, extending the applicability of online optimization to complex routing scenarios where reward structures are inherently non-linear and constrained.

Bandit FeedbackExpert RoutingMultinomial Subset Routing

This study addresses the limitations of flat routing in conventional Mixture-of-Experts (MoE) models, which often suffer from expert load imbalance and a lack of topological structure. To overcome these issues, this work proposes a hierarchical binary decision tree routing architecture that dynamically estimates branch probabilities via exponential moving averages to balance traffic across subtrees. This mechanism achieves load balancing without auxiliary losses while theoretically preventing routing collapse. Experimental results demonstrate that the proposed method preserves task accuracy across multiple benchmark datasets while significantly reducing cross-device communication overhead in distributed settings, thereby enabling efficient and balanced expert utilization.

Communication OverheadExpert Utilization ImbalanceFlat Routing

This study addresses the substantial memory overhead of deploying Mixture-of-Experts (MoE) models and the limitations of existing pruning methods, including poor alignment, high computational cost, and neglect of routing redundancy. We propose an efficient MoE pruning framework that achieves structured pruning by introducing learnable router biases and diversity regularization to precisely identify critical experts. Furthermore, an affine transformation-based expert approximation mechanism is designed to effectively compensate for accuracy degradation caused by pruning. Experimental results demonstrate that the proposed method successfully removes 25%–50% of experts across multiple large language models while consistently outperforming state-of-the-art algorithms on nine zero-shot benchmarks, achieving a favorable balance between model compression and reasoning capability.

expert rankingmemory reductionMixture-of-Experts

Hot Scholars

JW

Jingang Wang

Meituan
Information RetrievalNatural Language ProcessingMachine Translation
WC

William Chen

Carnegie Mellon University
Spoken Language ProcessingSpeech RecognitionSpeech TranslationMachine Translation
TQ

Tenghai Qiu

Institute of Automation,Chinese Academy of Sciences
intelligent decisiondeep reinforcement learningmulti-agent
XQ

Xiaojun Quan

Professor, School of Computer Science and Engineering, Sun Yat-sen University
natural language processingtext miningmachine learning