ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of effectively searching, distinguishing, and composing tools from large-scale tool libraries using large language models. To this end, it proposes a reinforcement learning-based framework for multi-turn search and fine-grained optimization. Methodologically, the framework introduces category-constrained tool discrimination, event-level search modeling, and trajectory-aligned credit assignment, integrated with agent reinforcement learning to enable iterative multi-turn search and fine-grained reward signal optimization. Experimental results demonstrate that the proposed approach significantly outperforms strong baselines on large-scale benchmarks, effectively enhancing both iterative search capabilities and complex tool composition performance.
📝 Abstract
Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively search, distinguish, and compose tools under context-length constraints. We identify large-scale tool selection as a new challenge for agentic reinforcement learning, highlighting that existing RL methods for knowledge-based question answering are inadequate for selecting tools while considering compatibility. To address this challenge, we propose ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection. Specifically, we introduce category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools, event-level search modeling to explicitly optimize the discovery of target tools during multi-turn search, and trajectory-aligned credit allocation to provide fine-grained reward signals for different stages of the search-selection process. Extensive experiments on large-scale tool selection benchmarks demonstrate that ToolSearcher consistently outperforms a set of strong baselines in challenging settings involving iterative search and complex tool composition.
Problem

Research questions and friction points this paper is trying to address.

Large-scale tool selection
Tool learning
Large language models
Reinforcement learning
Agentic AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Tool Selection
Large Language Models
Multi-turn Search
Credit Allocation
Z
Zhenlong Dai
Zhejiang University
X
Xujie Song
Zhejiang University
Z
Zitong Wang
Ant Group
T
Tong Niu
Ant Group
J
Jian Liu
Ant Group
W
Weiqiang Wang
Ant Group
X
Xiu Tang
Zhejiang University
Sai Wu
Sai Wu
Professor, Zhejiang University
Distributed DatabaseAI for DB
C
Chang Yao
Zhejiang University
Jingyuan Chen
Jingyuan Chen
Zhejiang University