Learning to Retrieve via Reinforcement Learning in Embedding Space

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional dense retrieval models struggle to directly optimize retrieval metrics and downstream task performance. To address this limitation, this work proposes RELER, a framework that leverages reinforcement learning to align task-specific rewards directly within the embedding space for retrieval optimization. The framework performs policy gradient updates via a REINFORCE algorithm grounded in the von Mises-Fisher distribution with an RLOO baseline, and introduces Conditional Mean Projection (CMP) to effectively suppress sampling noise during high-dimensional exploration. Experimental results demonstrate that RELER significantly outperforms InfoNCE-based methods on the BRIGHT benchmark, substantially improving both retrieval accuracy and generation quality in retrieval-augmented generation (RAG) scenarios.
📝 Abstract
Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards. We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO). As exploration in the high-dimensional embedding space is prone to sampling noise, we further propose conditional-mean projection (CMP), which projects each sampled embedding onto the low-dimensional subspace spanned by its encoder output and the candidate embeddings it is compared against, reducing noise in the policy gradient while preserving its expectation. We evaluate RELER on BRIGHT, a benchmark with reasoning-intensive queries that remain challenging for existing embedding models. RELER consistently outperforms InfoNCE and LambdaLoss in average nDCG@10 when post-training BGE-M3 and Qwen3-Embedding backbones. We further evaluate downstream utility through retrieval-augmented generation (RAG), where we adapt only the query encoder while keeping the document index and generator fixed. Across seven QA datasets, jointly optimizing retrieval and answer rewards improves both average retrieval performance and answer quality in RAG.
Problem

Research questions and friction points this paper is trying to address.

Dense Retrieval
Reinforcement Learning
Embedding Space
Retrieval-Augmented Generation
Reasoning-Intensive Queries
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Dense Retrieval
von Mises-Fisher Distribution
Conditional-Mean Projection
Retrieval-Augmented Generation
🔎 Similar Papers
2023-04-07Annual International ACM SIGIR Conference on Research and Development in Information RetrievalCitations: 21