🤖 AI Summary
This study addresses the challenges of credit assignment and low learning efficiency in reinforcement learning for search agents, which stem from reliance on sparse outcome rewards. Building upon large language models and Reinforcement Learning with Verifiable Rewards (RLVR), this work proposes a reward shaping mechanism that integrates intermediate retrieval signals with final outcomes. By systematically exploring strategies to optimize credit assignment, it constructs a multi-step trajectory training framework that jointly accounts for both the retrieval process and the final results. Experimental evaluations across multiple benchmarks demonstrate significant improvements in the overall performance of search agents. Furthermore, empirical analyses reveal the critical role of intermediate supervision signals in optimizing multi-step search behaviors, offering valuable insights for developing more effective and efficient reinforcement learning paradigms for complex search tasks.
📝 Abstract
Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.