PEARL: Plan Exploration and Adaptive Reinforcement Learning for Multihop Tool Use

📅 2026-01-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses key challenges in complex multi-hop tool use by large language models, including weak planning capabilities, tool hallucination, parameter errors, and poor interaction robustness. To overcome these limitations, the authors propose PEARL, a novel framework that uniquely integrates offline tool exploration with online reinforcement learning. During the offline phase, the model learns effective tool usage patterns and failure boundaries; in the online phase, a dedicated planner based on Group Relative Policy Optimization (GRPO) is employed, guided by a custom reward function designed to optimize planning quality. Evaluated on the ToolHop and T-Eval benchmarks, PEARL achieves state-of-the-art performance, attaining a 56.5% success rate on ToolHop while maintaining a low tool invocation error rate, thereby significantly enhancing both planning efficacy and robustness in multi-hop tool-augmented reasoning.

Technology Category

Planning, Routing, and Scheduling: Planning with Language ModelsMultiagent Systems: Multiagent PlanningHumans and AI: Human-Aware Planning and Behavior Prediction

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Large Language Models show great potential with external tools, but face significant challenges in complex, multi-turn tool invocation. They often exhibit weak planning, tool hallucination, erroneous parameter generation, and struggle with robust interaction. To tackle these issues, we present PEARL, a novel framework to enhance LLM planning and execution for sophisticated tool use. PEARL adopts a two-stage approach: an offline phase where the agent explores tools to learn valid usage patterns and failure conditions, and an online reinforcement learning phase. In the online phase, a dedicated Planner is trained via group Relative Policy Optimization (GRPO) with a carefully designed reward function that provides distinct signals for planning quality. Experiments on the ToolHop and T-Eval benchmarks show PEARL significantly outperforms existing methods, achieving a new state-of-the-art success rate of \textbf{56.5\%} on ToolHop while maintaining a low invocation error rate. Our work marks a key advance in addressing the complex planning challenges of tool use, contributing to the development of more robust and reliable LLM-based agents.
Problem

Research questions and friction points this paper is trying to address.

tool use
planning
hallucination
parameter generation
robust interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Plan Exploration
Adaptive Reinforcement Learning
Multihop Tool Use
Group Relative Policy Optimization
LLM-based Agents
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Q
Qihao Wang
Institute of Information Engineering, China; School of Cyber Security, University of Chinese Academy of Sciences, China
M
Mingzhe Lu
Institute of Information Engineering, China
J
Jiayue Wu
Institute of Information Engineering, China
Y
Yue Hu
Institute of Information Engineering, China
Y
Yanbing Liu
Institute of Information Engineering, China