Score
Designs and analyzes routing algorithms that select experts or routes by minimizing the combined encoding cost of model and data using Minimum Description Length principles, deriving routing objectives from description-length/MDL and formalizing model-complexity vs. performance trade-offs. Implements and evaluates selection and allocation schemes (including optimal expert counts per item/token), computes encoding-based routing objectives, and quantifies the complexity–performance tradeoffs implied by those objectives.
Large language model (LLM) systems commonly suffer from substantial resource waste due to static or suboptimal deployment strategies. Method: This paper proposes a cost-quality–aware query routing mechanism that dynamically dispatches user queries to the most suitable lightweight model, domain-specific expert, or embedding strategy. We formally define the routing problem for the first time and introduce a novel taxonomy jointly optimizing relevance and resource efficiency. Our framework systematically compares academic approaches with industrial practices, incorporating query understanding, policy selection, multi-granularity routing (at both model and embedding levels), fine-grained cost modeling, and a unified evaluation protocol. Contribution/Results: Experiments demonstrate that our mechanism significantly reduces inference overhead—by up to 42% in latency and 38% in GPU memory—while maintaining or even improving answer quality across diverse benchmarks. This work establishes both a theoretical foundation and a reproducible, practical paradigm for building efficient, scalable, and cost-effective LLM systems.
Static deployment of large language models (LLMs) struggles to dynamically select the optimal model based on query complexity and domain, leading to an imbalance between performance and cost. This work proposes a three-dimensional framework—centered on decision timing, information sources, and computation mechanisms—to systematically analyze dynamic routing and cascading strategies across multiple LLMs. It introduces the first taxonomy of routing paradigms spanning independently trained LLMs and integrates diverse technical approaches, including query difficulty estimation, human preference modeling, uncertainty quantification, reinforcement learning, and multimodal fusion. Experimental results demonstrate that well-designed routing systems can surpass the strongest individual model, achieving a superior trade-off between inference efficiency and performance, while also highlighting critical challenges such as generalization.
This work addresses the lack of effective methods to evaluate whether experts in sparse Mixture-of-Experts (MoE) models achieve non-redundant specialization, particularly in small-scale settings. To this end, the authors construct the first benchmark that accurately reflects large-scale routing behavior at a smaller scale, integrating multi-domain distinguishable data and an ideal reference router grounded in domain definitions. They systematically compare multiple routing strategies and introduce quantitative metrics to assess expert specialization. Their experiments reveal that a “balanced routing regime” is crucial for achieving both high expert utilization and meaningful specialization, a finding they further validate on models up to 35 times larger, demonstrating strong scalability of their conclusions.
This work addresses the limitations of static penalty strategies in VLSI global routing, which struggle to adapt to complex congestion topologies and jointly optimize congestion, wirelength, and via count. To overcome these challenges, the paper proposes a dynamic multi-objective optimization framework that reformulates rip-up-and-reroute (R&R) as a dynamic system. The approach integrates SHAP-driven congestion decomposition, 3D Dijkstra maze routing, and an adaptive PathFinder algorithm, and—novelty introduced here—employs a large language model as a semantic policy optimizer to dynamically tune penalty parameters under knowledge graph constraints. Evaluated on the ISPD 2025 benchmarks, the method reduces MEMPOOL overflow by 98.6%, lowers ARIANE overflow to 146,109 (a 29.8× improvement over the state of the art), and achieves a penalty score of 0.0538, significantly outperforming the prior best result of 1.780.
In sparsely activated Mixture-of-Experts (MoE) models, conventional top-k routing enforces fixed expert capacity constraints, leading to token dropping or inefficient padding and thus degrading hardware utilization; removing such constraints, however, causes severe load imbalance and reduced computational efficiency. To address this, we propose MaxScore routing—a novel mechanism that formulates token assignment as a minimum-cost maximum-flow problem. By integrating a differentiable SoftTopk operator with graph-based flow optimization, MaxScore achieves dynamic load balancing without explicit capacity limits, avoiding the limitations of iterative rerouting and optimal transport while preserving both differentiability and global optimality. Experiments demonstrate that, at identical FLOPs, models trained with MaxScore achieve lower training loss and higher downstream task performance, significantly improving computational efficiency and overall model effectiveness.
This work addresses a key limitation in sparse Mixture-of-Experts (MoE) models, where the routing mechanism jointly handles expert selection and output weighting, potentially constraining performance. The study provides the first systematic validation that these two functions should be decoupled and introduces Fixed Dispatch with Adaptive Aggregation (FDAA): a lightweight, learnable aggregation head is added atop a frozen backbone and fixed expert assignments, enabling end-to-end optimization of aggregation weights via the language modeling objective. Evaluated on pretrained MoE models such as OLMoE and DeepSeek-V2-Lite, FDAA achieves a 0.1523 reduction in cross-entropy on WikiText-103 and demonstrates consistent improvements across diverse benchmarks including C4 and PTB, confirming both the efficacy and generality of the proposed decoupling strategy.
This study addresses the challenge of optimizing non-additive rewards in online routing over expert subsets with bilateral constraints. We propose the Multi-Subset Routing (MSR) framework and an OMD-Approachability algorithm that integrates Online Mirror Descent with Blackwell’s approachability theory. This method overcomes the additive reward assumption inherent in traditional combinatorial bandits, achieving both constraint satisfaction and coverage maximization under winner-only feedback. Theoretical analysis demonstrates that both regret and constraint violation are bounded by O(1/√T). Furthermore, empirical evaluations on real-world crowdsourcing datasets validate the effectiveness of our approach. Collectively, this work establishes a novel paradigm for non-additive combinatorial online learning, extending the applicability of online optimization to complex routing scenarios where reward structures are inherently non-linear and constrained.
This study addresses the computational bottleneck of linear routers in large-scale Mixture-of-Experts (MoE) models by proposing the MoRE architecture, which leverages low-rank matrix decomposition to substantially reduce routing overhead and support larger expert counts. Theoretically, we prove that a logarithmic rank suffices to preserve routing expressivity and load balancing, thereby overcoming conventional full-rank constraints. From an engineering perspective, efficient deployment is achieved through Triton fused kernels and hardware-aware optimizations. Experimental results demonstrate that this approach maintains inference performance while effectively enhancing model memorization capacity and knowledge-based question answering capabilities, ultimately enabling significant scaling of expert numbers.
本文通过局部聚合视角分析混合专家模型,探讨路由、稀疏激活和共享专家等设计选择的统计作用,分离出逼近误差、专家学习误差和路由器估计误差。
This study reveals the failure of dynamic LLM routers in cost optimization. Through benchmarking and Pareto efficiency analysis, we demonstrate that commercial routers underperform random routing due to misaligned standard objective functions and a "difficulty blind spot," while the assumption of requiring large model pools proves invalid. To address these issues, we propose a novel evaluation framework and a dual-model routing paradigm that circumvents prevalent design pitfalls. Experimental results indicate that existing complex routing systems are broadly inefficient, whereas our streamlined dual-model architecture outperforms mainstream commercial solutions. This work establishes principled model selection strategies and provides critical theoretical and practical guidance for the design of LLM routing systems.
This study addresses the limitations of flat routing in conventional Mixture-of-Experts (MoE) models, which often suffer from expert load imbalance and a lack of topological structure. To overcome these issues, this work proposes a hierarchical binary decision tree routing architecture that dynamically estimates branch probabilities via exponential moving averages to balance traffic across subtrees. This mechanism achieves load balancing without auxiliary losses while theoretically preventing routing collapse. Experimental results demonstrate that the proposed method preserves task accuracy across multiple benchmark datasets while significantly reducing cross-device communication overhead in distributed settings, thereby enabling efficient and balanced expert utilization.
本文提出Expert-Space Exploration Reinforcement Learning (ESRL)方法,通过探索Mixture-of-Experts模型的专家路由空间,增加rollout多样性并保持计算路径可靠性,从而提高模型性能。