🤖 AI Summary
This paper addresses the limited reasoning capability of large language models (LLMs) in long-horizon, multi-turn legal document search tasks. We propose the first end-to-end reinforcement learning (RL) framework tailored for training multi-turn search agents. Our method formalizes search as a multi-step decision process and employs long-horizon RL to jointly optimize tool invocation and reasoning planning on a 14B-parameter LLM. Key contributions include: (i) the first systematic application of RL to multi-turn retrieval agent training, and (ii) empirical validation that extending the decision horizon significantly improves performance on complex search tasks. Experiments on a legal document search benchmark demonstrate an 85% accuracy—surpassing the state-of-the-art LLM baseline (78%)—and overcoming the performance ceiling imposed by conventional prompt engineering approaches.
📝 Abstract
Large Language Model (LLM) agents can leverage multiple turns and tools to solve complex tasks, with prompt-based approaches achieving strong performance. This work demonstrates that Reinforcement Learning (RL) can push capabilities significantly further by learning from experience. Through experiments on a legal document search benchmark, we show that our RL-trained 14 Billion parameter model outperforms frontier class models (85% vs 78% accuracy). In addition, we explore turn-restricted regimes, during training and at test-time, that show these agents achieve better results if allowed to operate over longer multi-turn horizons.