Score
Designs, builds, and analyzes algorithms and systems that decide how to forward or assign incoming requests, queries, or tasks to services, servers, models, experts, or agents, with the goal of optimizing metrics such as latency, throughput, load balancing, cache effectiveness, and security. This includes work on routing logic and optimization methods for multi‑model and mixture‑of‑experts inference, cache‑ and latency‑aware policies, secure and real‑time routing, general traffic/service routing, and domain‑specific variants such as vehicle routing algorithms.
Large language model (LLM) systems commonly suffer from substantial resource waste due to static or suboptimal deployment strategies. Method: This paper proposes a cost-quality–aware query routing mechanism that dynamically dispatches user queries to the most suitable lightweight model, domain-specific expert, or embedding strategy. We formally define the routing problem for the first time and introduce a novel taxonomy jointly optimizing relevance and resource efficiency. Our framework systematically compares academic approaches with industrial practices, incorporating query understanding, policy selection, multi-granularity routing (at both model and embedding levels), fine-grained cost modeling, and a unified evaluation protocol. Contribution/Results: Experiments demonstrate that our mechanism significantly reduces inference overhead—by up to 42% in latency and 38% in GPU memory—while maintaining or even improving answer quality across diverse benchmarks. This work establishes both a theoretical foundation and a reproducible, practical paradigm for building efficient, scalable, and cost-effective LLM systems.
Static deployment of large language models (LLMs) struggles to dynamically select the optimal model based on query complexity and domain, leading to an imbalance between performance and cost. This work proposes a three-dimensional framework—centered on decision timing, information sources, and computation mechanisms—to systematically analyze dynamic routing and cascading strategies across multiple LLMs. It introduces the first taxonomy of routing paradigms spanning independently trained LLMs and integrates diverse technical approaches, including query difficulty estimation, human preference modeling, uncertainty quantification, reinforcement learning, and multimodal fusion. Experimental results demonstrate that well-designed routing systems can surpass the strongest individual model, achieving a superior trade-off between inference efficiency and performance, while also highlighting critical challenges such as generalization.
This paper addresses dynamic customer routing optimization in skill-based service queueing systems (e.g., cloud data centers), where static policies fail to adapt to workload fluctuations and heterogeneous agent skills. We propose a UCB-based reinforcement learning routing algorithm that jointly estimates environment dynamics, balances multiple objectives—namely, waiting time minimization and load balancing—and incorporates parameter sensitivity analysis. To accelerate convergence and enhance robustness, we innovatively integrate heuristic rules into the exploration mechanism. Extensive experiments driven by real-world operational data demonstrate that our algorithm significantly outperforms standard baselines in both efficiency and adaptability. It achieves rapid online learning and self-adjustment under varying traffic conditions, validating its practical feasibility and effectiveness for deployment in complex, large-scale service systems.
Current large language model (LLM) inference services rely on generic heuristic strategies that overlook the unique dynamic structure of LLMs in request routing, scheduling, and KV cache management, leading to unstable performance and a lack of theoretical guarantees. This work presents the first systematic integration of operations research and machine learning systems to formally model the distinctive characteristics of LLM inference. Building upon this foundation, we propose an algorithmic framework grounded in mathematical optimization, queueing theory, and cache policy modeling. Our approach matches or surpasses existing heuristics across diverse workloads while providing provable performance bounds and enhanced predictability, thereby establishing a theoretically principled paradigm for LLM serving.
This work addresses the challenge of real-time request routing in large language model serving, where incoming requests must be scheduled to decoding nodes under constraints on batch size and KV cache capacity. Existing heuristic approaches struggle to explicitly balance the trade-off between latency and throughput. To overcome this limitation, the paper introduces a novel framework that integrates multi-objective optimization with online linear programming. Request admission decisions are made by comparing SLO-weighted rewards against dual shadow prices, while a warm-started first-order projected gradient method efficiently tracks dynamic dual variables, enabling millisecond-scale, interpretable, and tunable scheduling. Experiments on the Vidur simulation platform demonstrate that the proposed approach consistently outperforms baseline methods across diverse SLO configurations, achieving comprehensive improvements in end-to-end latency, time-to-first-token, throughput, and tail performance.
This work addresses the optimization challenges in large language model (LLM) inference arising from the entanglement of diverse workloads, routing strategies, and computational resource pools. To this end, we propose the Workload-Router-Pool (WRP) triadic co-design architecture—the first unified, systematic framework that decouples and jointly models the interactions among these three dimensions. The framework integrates key techniques including signal-driven routing, context-length pooling, semantic caching, multimodal agent routing, heterogeneous GPU pooling, KV-cache topology optimization, and reinforcement learning–guided model selection. Building upon the vLLM semantic router series, it forms a scalable WRP interaction matrix and delineates 21 concrete research directions and open challenges, offering a comprehensive blueprint for efficient, secure, and adaptive LLM inference systems.
This work addresses the critical yet overlooked role of model capability profiling in large language model (LLM) routing. The authors propose RouteProfile, a framework that formulates LLM profiling as structured information integration over heterogeneous interaction histories. For the first time, it systematically decouples profile design from routing mechanisms and explores the profile design space across four dimensions: organizational form, representation type, aggregation depth, and learning configuration. Experiments demonstrate that structured profiles outperform flat ones, query-level signals surpass domain-level signals, and trainable structured profiles significantly enhance generalization to unseen models. Evaluations across three representative routers consistently validate the superiority of the proposed approach, establishing a foundation for fair comparison and principled design of LLM routing systems.
This work addresses the trade-off between performance and cost in multi-model programming scenarios by proposing ACRouter, an agent-based dynamic model routing framework that reframes model selection as a context-action-feedback (C-A-F) loop, moving beyond conventional static classification paradigms. ACRouter comprises a coordinator, a validator, and a memory module, which collectively accumulate execution experience online to actively bridge task information gaps. Evaluated on CodeRouterBench—a newly curated benchmark encompassing eight state-of-the-art large language models—ACRouter achieves the lowest cumulative regret on in-distribution tasks and demonstrates strong generalization to out-of-distribution agent programming tasks.
This work addresses the limitations of static penalty strategies in VLSI global routing, which struggle to adapt to complex congestion topologies and jointly optimize congestion, wirelength, and via count. To overcome these challenges, the paper proposes a dynamic multi-objective optimization framework that reformulates rip-up-and-reroute (R&R) as a dynamic system. The approach integrates SHAP-driven congestion decomposition, 3D Dijkstra maze routing, and an adaptive PathFinder algorithm, and—novelty introduced here—employs a large language model as a semantic policy optimizer to dynamically tune penalty parameters under knowledge graph constraints. Evaluated on the ISPD 2025 benchmarks, the method reduces MEMPOOL overflow by 98.6%, lowers ARIANE overflow to 146,109 (a 29.8× improvement over the state of the art), and achieves a penalty score of 0.0538, significantly outperforming the prior best result of 1.780.
This work addresses the lack of a reproducible evaluation framework for meta-decision strategies—such as task decomposition and tool invocation—in existing agent systems. We introduce MetaRoute-Bench, the first open benchmark enabling fine-grained analysis of meta-decision routing, comprising 180 synthetic tasks, 8 distinct strategies, and 30 random seeds per configuration. Evaluation employs offline seeded execution and multidimensional metrics—including success rate, cost, and latency—to ensure fair comparison. Experiments demonstrate that task-aware compositional strategies achieve a significantly higher success rate (79.4%) compared to static strategies (76.7%), single-step routing (67.4%), and direct answering (52.9%), with only marginal increases in cost (4.7%) and latency (6.4%). Ablation studies further confirm the critical contributions of compositional operations and verification mechanisms. Code and execution trajectories are publicly released.
This work addresses the absence of a unified framework for fairly comparing and efficiently deploying large language model (LLM) routing strategies under diverse query and budget constraints. We propose the first formalized LLM routing framework, modeling routing as a sequential decision process that integrates joint encoding of context and models, configurable scoring functions, flexible decision rules, and automated learning signals—supporting single-turn, multi-turn, and personalized routing. To facilitate research and development, we introduce xRouteBench, a multitask benchmark, and release LLMRouter, a modular infrastructure integrating over 16 routing methods. Experiments demonstrate that learned routers outperform the strongest fixed-model baseline by 14.6%, lightweight variants excel under stringent cost constraints, and user-conditioned routing significantly enhances personalization effectiveness.
This work addresses the lack of reproducible, user-preference-driven evaluation frameworks in existing large language model (LLM) routing systems. We propose the first evaluation paradigm specifically focused on the routing layer, introducing RouteJudge—an online pairwise preference assessment platform that anonymously collects user comparisons of responses generated under different routing strategies. By attributing user preferences directly to routing decisions and integrating multidimensional factors such as cost, latency, and task metadata, our approach enables nuanced analysis. Complementing this, we develop ORBIT, a modular open-source toolbox offering unified interfaces for query representation, routing algorithm implementation, budget-aware evaluation, and method comparison, thereby supporting end-to-end standardized research workflows. Together, RouteJudge and ORBIT establish an open, extensible, and preference-aware ecosystem for LLM routing evaluation.