Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models

๐Ÿ“… 2026-07-09
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the critical challenge of dynamically allocating a limited query budget between resampling and rerouting strategies to maximize the answer accuracy of large language models. Treating these two approaches as competing strategies sharing a common budget, the paper proposes RoRโ€”an online, budget-aware test-time model selection method that dynamically allocates resources based on the marginal gain in accuracy per unit cost. RoR leverages a diverse model pool, online estimation of marginal returns, and a label-agnostic consistency verifier. Evaluated across four heterogeneous benchmarks, it significantly outperforms existing baselines, particularly in settings with high inter-model diversity, and achieves state-of-the-art trade-offs along the costโ€“accuracy Pareto frontier.
๐Ÿ“ Abstract
Routing among large language models (LLMs) trades response quality against serving cost, motivated by the reported gap between deployed routers and a per-instance oracle. Recent analysis shows that test-time resampling can recover per-instance selection headroom that no single-commit router captures; however, that guarantee holds only under an idealized oracle equipped with correctness labels and an unconstrained budget, neither of which a deployed system has. To the best of our knowledge, no previous work treats resampling the committed model and rerouting to an alternative model as competing uses of a single per-query cost budget. Therefore, this work formulates budget-aware test-time model selection: given a per-query budget and an imperfect verifier, allocate each unit of budget between resampling and rerouting so that expected correctness is maximized. An online resample-or-reroute (RoR) allocation policy driven by estimated marginal correctness per unit cost is proposed, and its behavior is grounded in the recoverability asymmetry between selection and sampling. Replay experiments on newly regenerated multi-draw correctness tensors from an eleven-model open-weight pool over four benchmarks of differing difficulty show that the proposed RoR policy attains a favorable cost-quality Pareto front relative to single-route, one-commit-router, budget-aware best-of-K, cascade, and random-allocation baselines for the tested pools, with the largest gains on the most heterogeneous benchmark; an ablation further shows the gains are verifier-gated, shrinking as verifier quality degrades, and robustness replays under a provider price vector and a label-free agreement verifier delineate where the conclusions carry over.
Problem

Research questions and friction points this paper is trying to address.

test-time model selection
budget-aware
large language models
resampling
rerouting
Innovation

Methods, ideas, or system contributions that make the work stand out.

budget-aware model selection
test-time adaptation
resampling vs rerouting
marginal correctness per cost
large language model routing
๐Ÿ”Ž Similar Papers
2024-07-08arXiv.orgCitations: 1
T
Teng-Ruei Chen
Institute of Bioinformatics and Systems Biology, National Yang Ming Chiao Tung University, Hsinchu 300, Taiwan; also with Krixvon, Taipei 100, Taiwan