Score
Designs and evaluates methods that generate sequential compositions of tools or actions—plans that specify which tools to apply and in what order—to accomplish multi-step tasks. This includes planners and algorithms that search combinatorial plan spaces or learn policies via reinforcement learning to optimize long-horizon objectives (e.g., final task quality) and produce sequential tool-ordering strategies.
This paper addresses the lack of a unified evaluation standard for large language models’ (LLMs) planning capabilities. We propose the first six-dimensional evaluation framework—comprising completeness, executability, optimality, representational capacity, generalizability, and efficiency—integrating classical AI planning theory with contemporary LLM empirical research. Through systematic bibliometric and comparative analysis across diverse tasks (e.g., web navigation, travel planning, database querying) and model architectures, we construct the first structured capability map of mainstream LLM-based planners, precisely delineating methodological boundaries. Our contributions are threefold: (1) the first multidimensional, unified evaluation framework for LLM planning; (2) an extensible, benchmarked analytical protocol; and (3) identification of three critical future research directions. The work provides both theoretical foundations and practical guidelines for evaluating and enhancing planning capabilities in agentic AI systems. (149 words)
Despite growing interest in leveraging large language models (LLMs) for planning—requiring environmental understanding, logical reasoning, and sequential decision-making—there exists no systematic taxonomy or standardized evaluation framework. Method: This paper introduces the first unified classification scheme for LLM-based planning methods, categorizing existing approaches into three paradigms: external module augmentation, fine-tuning-driven methods, and search-oriented techniques. It further establishes a standardized evaluation framework encompassing benchmark tasks, multidimensional metrics, and empirical comparisons. Contribution/Results: Through comprehensive literature analysis, methodological abstraction, and cross-paradigm mechanistic synthesis, this work delivers the field’s first holistic survey. It clarifies the technical evolution trajectory, identifies core bottlenecks—including scalability, generalization, and causal reasoning—and proposes future directions such as trustworthy planning, embodied collaboration, and neuro-symbolic integration. The study provides an authoritative knowledge graph and methodological roadmap for advancing LLM-based planning research.
This work addresses key challenges in complex multi-hop tool use by large language models, including weak planning capabilities, tool hallucination, parameter errors, and poor interaction robustness. To overcome these limitations, the authors propose PEARL, a novel framework that uniquely integrates offline tool exploration with online reinforcement learning. During the offline phase, the model learns effective tool usage patterns and failure boundaries; in the online phase, a dedicated planner based on Group Relative Policy Optimization (GRPO) is employed, guided by a custom reward function designed to optimize planning quality. Evaluated on the ToolHop and T-Eval benchmarks, PEARL achieves state-of-the-art performance, attaining a 56.5% success rate on ToolHop while maintaining a low tool invocation error rate, thereby significantly enhancing both planning efficacy and robustness in multi-hop tool-augmented reasoning.
To address the instability in planning and redundant error correction during multi-tool invocation by large language models (LLMs), this paper proposes a toolkit framework based on functional clustering. The method automatically groups tools into high-level semantic abstractions—toolkits—thereby elevating planning granularity from individual tools to toolkit-level units. It introduces a hierarchical prompting mechanism and a toolkit-aware planning-execution-feedback loop, enabling semantically consistent fault-tolerant re-planning. This design significantly reduces error-correction overhead while enhancing planning robustness and execution efficiency. Empirical evaluation across multiple benchmarks demonstrates substantial improvements in tool-call success rates (pass/wins) for both GPT-4 and Claude 3, validating the framework’s effectiveness in improving planning stability and error-recovery efficiency.
This work addresses the bottleneck in hierarchical planning for continuous-state/action domains—namely, its reliance on manually engineered symbolic predicates for state abstraction. We propose the first fully automated predicate invention framework. Methodologically, it employs grammar-guided differentiable search over predicate sets, jointly optimizing high-level operators and low-level samplers via proxy objective optimization and demonstration-guided symbolic predicate learning—enabling end-to-end abstraction discovery and hierarchical planning co-training. Evaluated on four robotic planning benchmarks, our approach significantly outperforms six baselines, demonstrating strong generalization: it rapidly solves unseen tasks. Our core contributions are twofold: (1) the first fully automated, human-intervention-free predicate invention mechanism; and (2) empirical validation that automatically discovered predicates substantially improve both planning efficiency and cross-task generalization.
Existing LLM-based tool learning methods predominantly formulate multi-step tool invocation as a text generation task, relying on supervised fine-tuning and thus struggling with the dynamic decision-making complexity inherent in sequential tool use. This work proposes the first step-wise reinforcement learning framework that explicitly models tool calling as a serialized decision process. We introduce a step-level reward shaping mechanism that separately quantifies the success and task-relevant contribution of each individual tool call. Further, we integrate policy gradient optimization with LLM–tool interface alignment to enable fine-grained policy updates. Evaluated on multi-step tool-use benchmarks, our approach achieves substantial improvements: +18.7% in task completion rate and +22.3% in tool-call accuracy, while significantly enhancing cross-step decision robustness. This framework establishes a novel paradigm for advancing LLM-based embodied intelligence and complex, multi-stage task execution.
This work addresses the limitations of existing large language model (LLM)-driven approaches to automated heuristic design, which are constrained by fixed evolutionary rules and static prompting templates, hindering long-horizon reasoning and efficient evolution. The authors propose modeling heuristic generation as a sequential decision-making process over an entailment graph, introducing the entailment graph as a stateful memory structure that enables cross-generation information reuse and conflict avoidance. A multi-agent collaborative framework is developed, comprising a policy agent that plans evolutionary actions, a world model agent that simulates heuristic performance, and a critic agent that performs routing-based reflection, thereby transforming trial-and-error evolution into state-aware, planning-driven search. The method achieves significantly faster convergence and yields superior heuristics across multiple combinatorial optimization problems, demonstrates compatibility with diverse LLM backbones, and exhibits strong scalability.
This work addresses the challenges of large action spaces, high stochasticity, long decision horizons, and resource constraints in sequential stochastic combinatorial optimization by proposing a model-based hierarchical reinforcement learning framework. The approach integrates a world model formulated as a semi-Markov decision process (SMDP) with a latent-space tree-search planner, constructing temporal structures of abstract actions through multi-timescale objectives and jointly learning budget-aware policies conditioned on subgoals to enable context-sensitive resource allocation. A key innovation lies in incorporating adaptive temporal abstraction into hierarchical planning, allowing the latent dynamics to explicitly capture the effective duration of abstract actions. Experimental results demonstrate that the proposed method significantly outperforms strong existing baselines across multiple challenging benchmarks.
This study reevaluates the effectiveness of the large language model PlanGPT in automated planning tasks, critically examining the reliability of its original claim regarding plan coverage. Through systematic comparisons on standard planning benchmarks, the work assesses PlanGPT against classical planners and greedy search algorithms, using plan cost and generation time as primary evaluation metrics—criteria newly introduced into the PlanGPT evaluation framework. The results reveal that PlanGPT’s performance is comparable only to that of a greedy search strategy, demonstrating no significant advantage in either plan quality or computational efficiency. These findings challenge the purported practical utility of PlanGPT in automated planning and raise questions about its added value over simpler, well-established methods.
This work addresses the common issue in hybrid discrete–continuous planning where first-order trajectories generated by conventional methods often violate second-order dynamical constraints of robotic systems, rendering them infeasible for execution. To bridge the gap between high-level task planning and low-level physical execution, the authors propose a reinforcement learning–based trajectory refinement framework that explicitly embeds analytical second-order dynamics into a Markov decision process. This approach continuously optimizes first-order trajectories produced by a high-level hybrid planner while respecting constraints on time windows, velocity, and acceleration. By integrating reinforcement learning with explicit second-order dynamical modeling—a combination not previously explored—the method significantly enhances the physical feasibility and real-world executability of planned trajectories.
This study addresses the challenge of designing an optimal recommendation mechanism in a finite-horizon discrete-time dynamic system where a system designer cannot directly control the actions of two strategic agents. The designer aims to maximize their own objective by recommending actions based on shared historical information, while ensuring that the agents find it sequentially rational to follow these recommendations, thereby forming a sequential rationality equilibrium. To this end, the paper proposes a novel recommendation mechanism that explicitly satisfies sequential rationality constraints and develops a computationally tractable solution framework combining backward induction with linear programming. This approach achieves, for the first time, the optimization of the designer’s objective under strict sequential rationality conditions, demonstrating both the effectiveness and computational feasibility of the proposed mechanism.
Existing benchmarks struggle to evaluate large language models’ long-horizon planning capabilities under conditions of limited tool visibility and dynamic disruptions. This work proposes the first interactive evaluation benchmark featuring an optional blocking mechanism, encompassing 327 retail tasks and 1,665 tools. By dynamically blocking tool access to simulate scenarios of tool unavailability or failure, and integrating task flows derived from large-scale API call trajectories with intermediate evidence-based reasoning, the benchmark systematically assesses agents’ abilities in iterative retrieval and path replanning. Experimental results reveal that even the strongest model tested (GPT-5.4) achieves only 51.90% accuracy without blocking, which sharply declines to 11.36% under severe blocking, exposing significant deficiencies in implicit failure detection and recovery over extended planning horizons.