🤖 AI Summary
This study addresses the challenge of effectively searching, distinguishing, and composing tools from large-scale tool libraries using large language models. To this end, it proposes a reinforcement learning-based framework for multi-turn search and fine-grained optimization. Methodologically, the framework introduces category-constrained tool discrimination, event-level search modeling, and trajectory-aligned credit assignment, integrated with agent reinforcement learning to enable iterative multi-turn search and fine-grained reward signal optimization. Experimental results demonstrate that the proposed approach significantly outperforms strong baselines on large-scale benchmarks, effectively enhancing both iterative search capabilities and complex tool composition performance.
📝 Abstract
Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively search, distinguish, and compose tools under context-length constraints. We identify large-scale tool selection as a new challenge for agentic reinforcement learning, highlighting that existing RL methods for knowledge-based question answering are inadequate for selecting tools while considering compatibility. To address this challenge, we propose ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection. Specifically, we introduce category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools, event-level search modeling to explicitly optimize the discovery of target tools during multi-turn search, and trajectory-aligned credit allocation to provide fine-grained reward signals for different stages of the search-selection process. Extensive experiments on large-scale tool selection benchmarks demonstrate that ToolSearcher consistently outperforms a set of strong baselines in challenging settings involving iterative search and complex tool composition.