FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing LLM routing methods score models independently, overlooking complementarity and enforcing rigid budgets, which often leads to redundant selections and low success rates. This work proposes FlexRouter, a framework that reformulates routing as a coverage-oriented subset selection problem, explicitly modeling model complementarity to maximize the probability that at least one selected model succeeds. During training, it introduces an objective based on failure-set marginalization to adaptively determine subset size. At inference, it employs Determinantal Point Processes (DPPs) to jointly capture capability and redundancy, coupled with a greedy strategy based on marginal log-determinant gain. Evaluated on the RouterEval benchmark, the proposed approach achieves higher coverage, reduced redundancy, and flexible inference costs compared to existing baselines.
📝 Abstract
Existing Large Language Model (LLM) routing methods score LLMs independently to select top-$k$ models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the overall probability of success. To address this, we propose FlexRouter, a routing framework that explicitly models model complementarity. FlexRouter optimizes for \textit{answer coverage}, maximizing the probability that at least one selected model yields a correct response. This objective aligns with practical inference pipelines where multiple candidate outputs are generated and a downstream verifier or user selects the final one. We formulate routing as a coverage-oriented subset selection problem and model the routing policy using Determinantal Point Processes (DPPs), which naturally capture both model competence and redundancy. To directly optimize coverage without requiring a ground-truth target subset, we introduce a training objective based on marginalizing over failure sets. During inference, we employ a greedy strategy based on marginal log-determinant gains, enabling the router to adaptively determine subset sizes without a predefined budget. Extensive experiments on the large-scale RouterEval benchmark demonstrate that our proposed FlexRouter achieves higher coverage with lower redundancy across both in-domain and out-of-domain tasks than strong baselines while maintaining flexible inference cost.
Problem

Research questions and friction points this paper is trying to address.

LLM routing
model complementarity
answer coverage
redundancy
flexible budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Routing
Determinantal Point Processes
Answer Coverage
Model Complementarity
Subset Selection
🔎 Similar Papers
No similar papers found.