When to Retrieve During Reasoning: Adaptive Retrieval for Large Reasoning Models

📅 2026-04-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing retrieval-augmented generation (RAG) systems, which statically inject context prior to reasoning and thus struggle to support the dynamic evidence acquisition required for multi-step inference in large language models. To overcome this, the authors propose ReaLM-Retrieve, a framework that adaptively triggers retrieval based on step-level uncertainty estimation and employs reinforcement learning to optimize retrieval timing decisions. Additionally, a lightweight ensembling mechanism is introduced, reducing per-retrieval computational overhead by 3.2×. Evaluated on three multi-hop question answering benchmarks—including MuSiQue—the method achieves an average F1 improvement of 10.1% while reducing retrieval calls by 47%. On MuSiQue specifically, it attains 71.2% F1 with only 1.8 retrievals per question and achieves a Recall@5 of 81.3%.
📝 Abstract
Large reasoning models such as DeepSeek-R1 and OpenAI o1 generate extended chains of thought spanning thousands of tokens, yet their integration with retrieval-augmented generation (RAG) remains fundamentally misaligned. Current RAG systems optimize for providing context before reasoning begins, while reasoning models require evidence injection during multi-step inference chains. We introduce ReaLM-Retrieve, a reasoning-aware retrieval framework that addresses this mismatch through three key innovations: (1) a step-level uncertainty detector that identifies knowledge gaps at reasoning-step granularity rather than token or sentence level; (2) a retrieval intervention policy that learns when external evidence maximally benefits ongoing reasoning; and (3) an efficiency-optimized integration mechanism that reduces per-retrieval overhead by 3.2x compared to naive integration. Experiments on MuSiQue, HotpotQA, and 2WikiMultiHopQA demonstrate that ReaLM-Retrieve achieves on average 10.1% absolute improvement in answer F1 over standard RAG (range: 9.0-11.8% across the three benchmarks) while reducing retrieval calls by 47% compared to fixed-interval approaches like IRCoT (all improvements significant at p<0.01, paired bootstrap). On the challenging MuSiQue benchmark requiring 2-4 hop reasoning, our method achieves 71.2% F1 with an average of only 1.8 retrieval calls per question. Analysis shows that ReaLM-Retrieve also improves retrieval quality itself, achieving 81.3% Recall@5 with consistently higher precision and MRR than fixed-interval baselines on supporting evidence, establishing new state-of-the-art efficiency-accuracy trade-offs for reasoning-intensive retrieval tasks.
Problem

Research questions and friction points this paper is trying to address.

retrieval-augmented generation
reasoning models
adaptive retrieval
multi-step inference
knowledge gaps
Innovation

Methods, ideas, or system contributions that make the work stand out.

adaptive retrieval
reasoning-aware RAG
step-level uncertainty detection
retrieval intervention policy
efficient retrieval integration
D
Dongxin Guo
The University of Hong Kong; Brain Investing Limited
J
Jikun Wu
Stellaris AI Limited; Brain Investing Limited
S
Siu Ming Yiu
The University of Hong Kong; Brain Investing Limited