Reinforcement Learning for Long-Horizon Multi-Turn Search Agents

📅 2025-10-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the limited reasoning capability of large language models (LLMs) in long-horizon, multi-turn legal document search tasks. We propose the first end-to-end reinforcement learning (RL) framework tailored for training multi-turn search agents. Our method formalizes search as a multi-step decision process and employs long-horizon RL to jointly optimize tool invocation and reasoning planning on a 14B-parameter LLM. Key contributions include: (i) the first systematic application of RL to multi-turn retrieval agent training, and (ii) empirical validation that extending the decision horizon significantly improves performance on complex search tasks. Experiments on a legal document search benchmark demonstrate an 85% accuracy—surpassing the state-of-the-art LLM baseline (78%)—and overcoming the performance ceiling imposed by conventional prompt engineering approaches.

Technology Category

Search and Optimization: Learning to SearchMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language Models

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Large Language Model (LLM) agents can leverage multiple turns and tools to solve complex tasks, with prompt-based approaches achieving strong performance. This work demonstrates that Reinforcement Learning (RL) can push capabilities significantly further by learning from experience. Through experiments on a legal document search benchmark, we show that our RL-trained 14 Billion parameter model outperforms frontier class models (85% vs 78% accuracy). In addition, we explore turn-restricted regimes, during training and at test-time, that show these agents achieve better results if allowed to operate over longer multi-turn horizons.
Problem

Research questions and friction points this paper is trying to address.

Enhancing multi-turn search agents using reinforcement learning
Improving legal document search accuracy through RL training
Exploring longer turn horizons for better agent performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses reinforcement learning for multi-turn search
Trains 14B parameter model on legal documents
Extends agent operation over longer horizons
🔎 Similar Papers
No similar papers found.