mdl-based routing

Designs and analyzes routing algorithms that select experts or routes by minimizing the combined encoding cost of model and data using Minimum Description Length principles, deriving routing objectives from description-length/MDL and formalizing model-complexity vs. performance trade-offs. Implements and evaluates selection and allocation schemes (including optimal expert counts per item/token), computes encoding-based routing objectives, and quantifies the complexity–performance tradeoffs implied by those objectives.

mdl-basedrouting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.43
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Static deployment of large language models (LLMs) struggles to dynamically select the optimal model based on query complexity and domain, leading to an imbalance between performance and cost. This work proposes a three-dimensional framework—centered on decision timing, information sources, and computation mechanisms—to systematically analyze dynamic routing and cascading strategies across multiple LLMs. It introduces the first taxonomy of routing paradigms spanning independently trained LLMs and integrates diverse technical approaches, including query difficulty estimation, human preference modeling, uncertainty quantification, reinforcement learning, and multimodal fusion. Experimental results demonstrate that well-designed routing systems can surpass the strongest individual model, achieving a superior trade-off between inference efficiency and performance, while also highlighting critical challenges such as generalization.

dynamic model routingLLM inferencemodel selection

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of effective methods to evaluate whether experts in sparse Mixture-of-Experts (MoE) models achieve non-redundant specialization, particularly in small-scale settings. To this end, the authors construct the first benchmark that accurately reflects large-scale routing behavior at a smaller scale, integrating multi-domain distinguishable data and an ideal reference router grounded in domain definitions. They systematically compare multiple routing strategies and introduce quantitative metrics to assess expert specialization. Their experiments reveal that a “balanced routing regime” is crucial for achieving both high expert utilization and meaningful specialization, a finding they further validate on models up to 35 times larger, demonstrating strong scalability of their conclusions.

expert specializationlarge language modelsMixture-of-Experts

This work addresses the limitations of static penalty strategies in VLSI global routing, which struggle to adapt to complex congestion topologies and jointly optimize congestion, wirelength, and via count. To overcome these challenges, the paper proposes a dynamic multi-objective optimization framework that reformulates rip-up-and-reroute (R&R) as a dynamic system. The approach integrates SHAP-driven congestion decomposition, 3D Dijkstra maze routing, and an adaptive PathFinder algorithm, and—novelty introduced here—employs a large language model as a semantic policy optimizer to dynamically tune penalty parameters under knowledge graph constraints. Evaluated on the ISPD 2025 benchmarks, the method reduces MEMPOOL overflow by 98.6%, lowers ARIANE overflow to 146,109 (a 29.8× improvement over the state of the art), and achieves a penalty score of 0.0538, significantly outperforming the prior best result of 1.780.

congestion minimizationmulti-objective optimizationNP-hard combinatorial optimization

Maximum Score Routing For Mixture-of-Experts

Aug 18, 2025
BD
Bowen Dong
🏛️ Tsinghua University | Tianjin University | Seed-Foundation-Model Team | ByteDance

In sparsely activated Mixture-of-Experts (MoE) models, conventional top-k routing enforces fixed expert capacity constraints, leading to token dropping or inefficient padding and thus degrading hardware utilization; removing such constraints, however, causes severe load imbalance and reduced computational efficiency. To address this, we propose MaxScore routing—a novel mechanism that formulates token assignment as a minimum-cost maximum-flow problem. By integrating a differentiable SoftTopk operator with graph-based flow optimization, MaxScore achieves dynamic load balancing without explicit capacity limits, avoiding the limitations of iterative rerouting and optimal transport while preserving both differentiability and global optimality. Experiments demonstrate that, at identical FLOPs, models trained with MaxScore achieve lower training loss and higher downstream task performance, significantly improving computational efficiency and overall model effectiveness.

Addresses token dropping in MoE due to expert capacity constraintsBalances load and computation without iterative rerouting limitationsImproves hardware efficiency by reducing padding in underutilized experts

This work addresses a key limitation in sparse Mixture-of-Experts (MoE) models, where the routing mechanism jointly handles expert selection and output weighting, potentially constraining performance. The study provides the first systematic validation that these two functions should be decoupled and introduces Fixed Dispatch with Adaptive Aggregation (FDAA): a lightweight, learnable aggregation head is added atop a frozen backbone and fixed expert assignments, enabling end-to-end optimization of aggregation weights via the language modeling objective. Evaluated on pretrained MoE models such as OLMoE and DeepSeek-V2-Lite, FDAA achieves a 0.1523 reduction in cross-entropy on WikiText-103 and demonstrates consistent improvements across diverse benchmarks including C4 and PTB, confirming both the efficacy and generality of the proposed decoupling strategy.

expert aggregationexpert dispatchMixture-of-Experts

This study addresses the challenge of optimizing non-additive rewards in online routing over expert subsets with bilateral constraints. We propose the Multi-Subset Routing (MSR) framework and an OMD-Approachability algorithm that integrates Online Mirror Descent with Blackwell’s approachability theory. This method overcomes the additive reward assumption inherent in traditional combinatorial bandits, achieving both constraint satisfaction and coverage maximization under winner-only feedback. Theoretical analysis demonstrates that both regret and constraint violation are bounded by O(1/√T). Furthermore, empirical evaluations on real-world crowdsourcing datasets validate the effectiveness of our approach. Collectively, this work establishes a novel paradigm for non-additive combinatorial online learning, extending the applicability of online optimization to complex routing scenarios where reward structures are inherently non-linear and constrained.

Bandit FeedbackExpert RoutingMultinomial Subset Routing

Latest Papers

What's happening recently
View more

This study addresses the computational bottleneck of linear routers in large-scale Mixture-of-Experts (MoE) models by proposing the MoRE architecture, which leverages low-rank matrix decomposition to substantially reduce routing overhead and support larger expert counts. Theoretically, we prove that a logarithmic rank suffices to preserve routing expressivity and load balancing, thereby overcoming conventional full-rank constraints. From an engineering perspective, efficient deployment is achieved through Triton fused kernels and hardware-aware optimizations. Experimental results demonstrate that this approach maintains inference performance while effectively enhancing model memorization capacity and knowledge-based question answering capabilities, ultimately enabling significant scaling of expert numbers.

hardware-aware inferencelow-rank routingMixture-of-Experts

This study reveals the failure of dynamic LLM routers in cost optimization. Through benchmarking and Pareto efficiency analysis, we demonstrate that commercial routers underperform random routing due to misaligned standard objective functions and a "difficulty blind spot," while the assumption of requiring large model pools proves invalid. To address these issues, we propose a novel evaluation framework and a dual-model routing paradigm that circumvents prevalent design pitfalls. Experimental results indicate that existing complex routing systems are broadly inefficient, whereas our streamlined dual-model architecture outperforms mainstream commercial solutions. This work establishes principled model selection strategies and provides critical theoretical and practical guidance for the design of LLM routing systems.

Dynamic LLM RoutersEvaluation MethodologyInference Cost

This study addresses the limitations of flat routing in conventional Mixture-of-Experts (MoE) models, which often suffer from expert load imbalance and a lack of topological structure. To overcome these issues, this work proposes a hierarchical binary decision tree routing architecture that dynamically estimates branch probabilities via exponential moving averages to balance traffic across subtrees. This mechanism achieves load balancing without auxiliary losses while theoretically preventing routing collapse. Experimental results demonstrate that the proposed method preserves task accuracy across multiple benchmark datasets while significantly reducing cross-device communication overhead in distributed settings, thereby enabling efficient and balanced expert utilization.

Communication OverheadExpert Utilization ImbalanceFlat Routing

本文提出Expert-Space Exploration Reinforcement Learning (ESRL)方法,通过探索Mixture-of-Experts模型的专家路由空间,增加rollout多样性并保持计算路径可靠性,从而提高模型性能。

Expert RoutingMixture-of-ExpertsReinforcement Learning

Hot Scholars

ES

Emmanuel S. Pilli

Associate Professor, Department of CSE, Malaviya National Institute of Technology Jaipur
SecurityForensicsCloud ComputingBig Data