π€ AI Summary
This study addresses the high inference costs of vision-language model reranking in multimodal retrieval and the resource inefficiency caused by fixed scheduling. To this end, it proposes an Adaptive Budget Elo Reranking framework that models local outputs as tournaments and maintains a global state through continuous Elo rating updates, enabling bi-level dynamic computation allocation. The method further optimizes global resources by integrating an online budgeting mechanism with submodular information-theoretic motivation and Bradley-Terry log-likelihood gradient ascent. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance on benchmarks such as CIRR under equivalent invocation budgets, while maintaining high efficiency even in low-budget scenarios.
π Abstract
Vision-language models (VLMs) are powerful listwise rerankers for multimodal retrieval, but high inference costs restrict them to evaluating small local candidate views. Existing multi-call strategies rely on fixed schedules, wasting expensive VLM calls on uninformative candidate pairs and easy queries. To address this, we propose Adaptive Multi-view Budgeted Elo Reranking (AMBER), an online, budgeted multi-view reranking framework that dynamically optimizes global resource allocation. AMBER treats fragmented listwise VLM outputs as local tournaments, using continuous Elo updates to maintain a lightweight global ranking state. Building on this, it allocates computation at two levels: dynamically constructing candidate views with high score ambiguity, and scheduling queries to maximize expected information gain. We show that each Elo update corresponds to a stochastic gradient ascent step on the Bradley-Terry log-likelihood, and provide a submodular information-theoretic motivation for the query-level allocation strategy. Experiments on CIRR, CIRCO, and PhotoBench demonstrate that AMBER achieves the strongest overall performance among the compared multi-call VLM reranking methods under comparable VLM-call budgets, while remaining effective in lower-budget settings. Our code is publicly available at https://github.com/wnlfc/AMBER.