The Routing Plateau: Understanding the Accuracy Limits of LLM Routers

📅 2026-05-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the “routing plateau” phenomenon in existing LLM routing methods, where accuracy converges within a narrow range significantly below the theoretical optimum. We propose the “correctness prediction bottleneck” hypothesis and construct Nine-by-300k, a large-scale benchmark for systematically evaluating diverse routing paradigms, including clustering, classifiers, ranking, and confidence-based approaches, alongside comparative analyses incorporating end-to-end fine-tuning and data scaling. Our findings confirm that current routers capture only global average trends while lacking instance-level fine-grained signals. Furthermore, merely increasing data volume or optimizing architectures fails to overcome this performance ceiling, yielding marginal improvements of only 1.24 percentage points. Consequently, we identify the incorporation of novel signals, such as partial output trajectories, as a critical direction for achieving future breakthroughs in LLM routing performance.
📝 Abstract
LLM routing has become a popular approach to improve the cost-quality trade-off of LLM services by adaptively selecting a model for each query. Recent work has explored a broad range of routing methods, including clustering-based routers, learned classifiers, pairwise ranking, and confidence-based approaches. Our extensive study of 21 routing methods across five benchmarks reveals a consistent phenomenon that we call the routing plateau (Fig. 1): many methods, including kNN, achieve very similar accuracy and converge to a narrow performance range that remains far below the oracle router. Our analysis supports a correctness-prediction bottleneck hypothesis: current routers primarily learn global-average model performance trends rather than fine-grained, query-specific routing signals. As a result, they collectively fail on queries that require instance-specific routing decisions. Moreover, to understand whether the plateau can be alleviated with a better training setup, we construct a 300K-query benchmark (Nine-by-300k). More data, larger encoders, and end-to-end fine-tuning improve eight routers by 1.24 pp on average, but leave the plateau largely intact. These findings suggest that further progress may require inputs beyond the query itself, such as partial output trajectories that reveal how models attempt the task.
Problem

Research questions and friction points this paper is trying to address.

LLM routing
routing plateau
accuracy limits
correctness-prediction bottleneck
query-specific routing
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Routing
Routing Plateau
Correctness-Prediction Bottleneck
Benchmark Construction
Cost-Quality Trade-off
🔎 Similar Papers
No similar papers found.
Y
Yifan Lu
Rice University
Q
Qiyue Zhang
Rice University
S
Shenrun Zhang
Rice University
Z
Zhibo Yu
Rice University
Z
Zhuang Wang
Amazon
Hanjie Chen
Hanjie Chen
Rice University
Natural Language ProcessingInterpretable Machine Learning
Jiarong Xing
Jiarong Xing
UC Berkeley; Rice University
SystemsNetworkingSecurity