Dynamic LLM Routers are Often Misguided

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study reveals the failure of dynamic LLM routers in cost optimization. Through benchmarking and Pareto efficiency analysis, we demonstrate that commercial routers underperform random routing due to misaligned standard objective functions and a "difficulty blind spot," while the assumption of requiring large model pools proves invalid. To address these issues, we propose a novel evaluation framework and a dual-model routing paradigm that circumvents prevalent design pitfalls. Experimental results indicate that existing complex routing systems are broadly inefficient, whereas our streamlined dual-model architecture outperforms mainstream commercial solutions. This work establishes principled model selection strategies and provides critical theoretical and practical guidance for the design of LLM routing systems.
📝 Abstract
Dynamic LLM routers promise to cut inference costs by sending each query to the cheapest model that can answer it correctly. We analyze six commercial routers across 14 settings on a diverse benchmark spanning eight task categories, finding that none of them outperforms a router that randomly selects between two well-chosen models at matched cost. Some underperform by more than 10 percentage points. We trace this gap to four patterns prevalent across routers: difficulty blindness, length reversal, semantic matching, and roster suboptimality. We show that the first three are what the standard objective rewards: cost-accuracy Pareto efficiency on realized costs favors escalating moderately hard queries over the hardest ones, shorter queries over longer ones, and routing by a query's source over its difficulty. We also argue that the two assumptions that would justify large rosters, model granularity and model specialization, do not hold empirically. We propose an alternative evaluation methodology that does not reward these patterns, and as a proof of concept, we design a simple two-model router that avoids all four. Nevertheless, its gain over random routing is limited, because a well-chosen roster leaves little to route.
Problem

Research questions and friction points this paper is trying to address.

Dynamic LLM Routers
Inference Cost
Routing Performance
Model Selection
Evaluation Methodology
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic LLM Router
Evaluation Methodology
Pareto Efficiency
Routing Patterns
Model Roster
Sam Wang
Sam Wang
Professor of Neuroscience, Princeton University
NeuroscienceStatistical PoliticsTwo-photon microscopyAutismCerebellum
J
Julia White
Fastino Labs
S
Sahibzada Allahyar
Fastino Labs
D
Dhruv Atreja
Fastino Labs
U
Urchade Zaratiana
Fastino Labs
K
Kelton Zhang
Fastino Labs