🤖 AI Summary
This study addresses the “routing plateau” phenomenon in existing LLM routing methods, where accuracy converges within a narrow range significantly below the theoretical optimum. We propose the “correctness prediction bottleneck” hypothesis and construct Nine-by-300k, a large-scale benchmark for systematically evaluating diverse routing paradigms, including clustering, classifiers, ranking, and confidence-based approaches, alongside comparative analyses incorporating end-to-end fine-tuning and data scaling. Our findings confirm that current routers capture only global average trends while lacking instance-level fine-grained signals. Furthermore, merely increasing data volume or optimizing architectures fails to overcome this performance ceiling, yielding marginal improvements of only 1.24 percentage points. Consequently, we identify the incorporation of novel signals, such as partial output trajectories, as a critical direction for achieving future breakthroughs in LLM routing performance.
📝 Abstract
LLM routing has become a popular approach to improve the cost-quality trade-off of LLM services by adaptively selecting a model for each query. Recent work has explored a broad range of routing methods, including clustering-based routers, learned classifiers, pairwise ranking, and confidence-based approaches. Our extensive study of 21 routing methods across five benchmarks reveals a consistent phenomenon that we call the routing plateau (Fig. 1): many methods, including kNN, achieve very similar accuracy and converge to a narrow performance range that remains far below the oracle router. Our analysis supports a correctness-prediction bottleneck hypothesis: current routers primarily learn global-average model performance trends rather than fine-grained, query-specific routing signals. As a result, they collectively fail on queries that require instance-specific routing decisions. Moreover, to understand whether the plateau can be alleviated with a better training setup, we construct a 300K-query benchmark (Nine-by-300k). More data, larger encoders, and end-to-end fine-tuning improve eight routers by 1.24 pp on average, but leave the plateau largely intact. These findings suggest that further progress may require inputs beyond the query itself, such as partial output trajectories that reveal how models attempt the task.