🤖 AI Summary
This study addresses the limitation of fixed query representations in general-purpose multimodal retrieval, which hinders the exploitation of retrieved results to clarify information needs. To this end, we propose the Alternating Retrieval and Refinement (ARR) model, which introduces an autoregressive feedback mechanism that dynamically optimizes information need representations by alternating between retrieval and query embedding updates. During training, ARR combines stepwise contrastive supervised fine-tuning with reinforcement learning to precisely select informative feedback items, while employing a query-side adapter to enhance generalization. Experiments demonstrate that ARR significantly outperforms existing baselines on both in-domain and zero-shot benchmarks, validating that iterative feedback yields dual benefits for initial retrieval accuracy and downstream reasoning performance.
📝 Abstract
Universal multimodal retrieval typically encodes a query once and ranks independently indexed items by embedding similarity. This design supports efficient search, but leaves the query representation unchanged even when retrieved items could help clarify the information need. We introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that learns both to select informative items and to use their content to refine subsequent retrieval. ARR alternates between retrieving an item and updating the query embedding, then uses the final embedding to rank the collection. Supervised fine-tuning teaches the encoder to use feedback through stepwise contrastive supervision. Reinforcement learning treats feedback items as actions and optimizes their selection using the final reciprocal rank of a relevant item. A query-side adapter enables this optimization against a fixed item index. ARR demonstrates strong retrieval performance on both in-domain and zero-shot benchmarks, outperforming the compared baselines on average. Further analyses show that feedback improves retrieval at inference time and that training with feedback also improves the initial query embedding, before any item is observed.