🤖 AI Summary
This work addresses key challenges in complex multi-hop tool use by large language models, including weak planning capabilities, tool hallucination, parameter errors, and poor interaction robustness. To overcome these limitations, the authors propose PEARL, a novel framework that uniquely integrates offline tool exploration with online reinforcement learning. During the offline phase, the model learns effective tool usage patterns and failure boundaries; in the online phase, a dedicated planner based on Group Relative Policy Optimization (GRPO) is employed, guided by a custom reward function designed to optimize planning quality. Evaluated on the ToolHop and T-Eval benchmarks, PEARL achieves state-of-the-art performance, attaining a 56.5% success rate on ToolHop while maintaining a low tool invocation error rate, thereby significantly enhancing both planning efficacy and robustness in multi-hop tool-augmented reasoning.
📝 Abstract
Large Language Models show great potential with external tools, but face significant challenges in complex, multi-turn tool invocation. They often exhibit weak planning, tool hallucination, erroneous parameter generation, and struggle with robust interaction. To tackle these issues, we present PEARL, a novel framework to enhance LLM planning and execution for sophisticated tool use. PEARL adopts a two-stage approach: an offline phase where the agent explores tools to learn valid usage patterns and failure conditions, and an online reinforcement learning phase. In the online phase, a dedicated Planner is trained via group Relative Policy Optimization (GRPO) with a carefully designed reward function that provides distinct signals for planning quality. Experiments on the ToolHop and T-Eval benchmarks show PEARL significantly outperforms existing methods, achieving a new state-of-the-art success rate of \textbf{56.5\%} on ToolHop while maintaining a low invocation error rate. Our work marks a key advance in addressing the complex planning challenges of tool use, contributing to the development of more robust and reliable LLM-based agents.