🤖 AI Summary
Existing LLM routing methods score models independently, overlooking complementarity and enforcing rigid budgets, which often leads to redundant selections and low success rates. This work proposes FlexRouter, a framework that reformulates routing as a coverage-oriented subset selection problem, explicitly modeling model complementarity to maximize the probability that at least one selected model succeeds. During training, it introduces an objective based on failure-set marginalization to adaptively determine subset size. At inference, it employs Determinantal Point Processes (DPPs) to jointly capture capability and redundancy, coupled with a greedy strategy based on marginal log-determinant gain. Evaluated on the RouterEval benchmark, the proposed approach achieves higher coverage, reduced redundancy, and flexible inference costs compared to existing baselines.
📝 Abstract
Existing Large Language Model (LLM) routing methods score LLMs independently to select top-$k$ models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the overall probability of success. To address this, we propose FlexRouter, a routing framework that explicitly models model complementarity. FlexRouter optimizes for \textit{answer coverage}, maximizing the probability that at least one selected model yields a correct response. This objective aligns with practical inference pipelines where multiple candidate outputs are generated and a downstream verifier or user selects the final one. We formulate routing as a coverage-oriented subset selection problem and model the routing policy using Determinantal Point Processes (DPPs), which naturally capture both model competence and redundancy. To directly optimize coverage without requiring a ground-truth target subset, we introduce a training objective based on marginalizing over failure sets. During inference, we employ a greedy strategy based on marginal log-determinant gains, enabling the router to adaptively determine subset sizes without a predefined budget. Extensive experiments on the large-scale RouterEval benchmark demonstrate that our proposed FlexRouter achieves higher coverage with lower redundancy across both in-domain and out-of-domain tasks than strong baselines while maintaining flexible inference cost.