Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval
This study addresses the limitation of fixed query representations in general-purpose multimodal retrieval, which hinders the exploitation of retrieved results to clarify information needs. To this end, we propose the Alternating Retrieval and Refinement (ARR) model, which introduces an autoregressive feedback mechanism that dynamically optimizes information need representations by alternating between retrieval and query embedding updates. During training, ARR combines stepwise contrastive supervised fine-tuning with reinforcement learning to precisely select informative feedback items, while employing a query-side adapter to enhance generalization. Experiments demonstrate that ARR significantly outperforms existing baselines on both in-domain and zero-shot benchmarks, validating that iterative feedback yields dual benefits for initial retrieval accuracy and downstream reasoning performance.