Score
Designs and builds methods that use large language models to decompose complex tasks into hierarchical, span- and step-aware plans and to synthesize executable subgoal sequences, persona-conditioned agent simulations, and alternative plan variants for decision making. Analyzes and implements LLM planners, prompting strategies, and learning-augmented optimization techniques that enable LLM-driven agents to generate, adapt, evaluate, and simulate long‑horizon plans and trajectories.
Despite growing interest in leveraging large language models (LLMs) for planning—requiring environmental understanding, logical reasoning, and sequential decision-making—there exists no systematic taxonomy or standardized evaluation framework. Method: This paper introduces the first unified classification scheme for LLM-based planning methods, categorizing existing approaches into three paradigms: external module augmentation, fine-tuning-driven methods, and search-oriented techniques. It further establishes a standardized evaluation framework encompassing benchmark tasks, multidimensional metrics, and empirical comparisons. Contribution/Results: Through comprehensive literature analysis, methodological abstraction, and cross-paradigm mechanistic synthesis, this work delivers the field’s first holistic survey. It clarifies the technical evolution trajectory, identifies core bottlenecks—including scalability, generalization, and causal reasoning—and proposes future directions such as trustworthy planning, embodied collaboration, and neuro-symbolic integration. The study provides an authoritative knowledge graph and methodological roadmap for advancing LLM-based planning research.
This work investigates the fundamental applicability boundaries of large language models (LLMs) in automated planning. Through a systematic literature review, a multi-dimensional capability assessment framework, and empirical evaluation on canonical domains—including Block World and Logistics—the study reveals critical limitations: inconsistent long-horizon reasoning, failure in constraint-sensitive planning, and unreliable state tracking. Methodologically, it employs rigorous comparative analysis across diverse planning tasks to isolate intrinsic LLM deficiencies. The primary contribution is the first principled argument that LLMs are unsuitable as standalone planners; instead, it proposes “hybrid intelligent planning”—a novel paradigm wherein LLMs serve exclusively as semantic understanding and heuristic generation modules, tightly integrated with symbolic reasoning engines and search algorithms. The work establishes a reproducible, taxonomy-based evaluation methodology and provides concrete architectural design principles for synergistic LLM–symbolic system integration, thereby delivering both theoretical foundations and practical guidelines for LLM-augmented planning.
Large language models (LLMs) exhibit limited planning capabilities in multi-step reasoning and goal-directed tasks. To address this, we propose the Modular Agent Planning (MAP) architecture—a cognitively inspired, reinforcement learning–informed framework that decomposes planning into specialized LLM modules: conflict monitoring, state prediction, and task decomposition. These modules operate in a recurrent, collaborative loop to enable dynamic, adaptive planning. MAP supports lightweight deployment and cross-task generalization without fine-tuning, seamlessly adapting to diverse LLM scales (e.g., Llama3-70B). Empirical evaluation on graph traversal, Tower of Hanoi, PlanBench, and StrategyQA demonstrates substantial improvements over zero-shot prompting, chain-of-thought, and tree-of-thought baselines—yielding higher planning accuracy and robustness. Our core contribution is the first systematic integration of modular, division-of-labor mechanisms into LLM-based planning, establishing a novel paradigm for structured, interpretable, and scalable agent reasoning.
Traditional large language models (LLMs) suffer from low reliability, poor completeness, and high reasoning overhead when directly generating planning solutions, limiting their generalization to unseen tasks. This work proposes a novel paradigm that leverages LLMs during the construction phase to automatically generate maintainable and verifiable symbolic planners, while entirely eliminating reliance on LLMs during the inference phase—thereby achieving efficient, lightweight, and reliable planning. The study systematically formulates methods for LLM-driven planner generation, delineates three distinct evolutionary pathways, and establishes a theoretical framework to guide the development of sustainable and verifiable agent planning systems.
This paper addresses the limitations of large language model (LLM)-based agents in long-horizon planning, dynamic interaction, and complex decision-making within intricate environments. Methodologically, it introduces the first systematic optimization survey, proposing a unified classification framework that dichotomizes optimization strategies into parameter-driven approaches (e.g., supervised fine-tuning, PPO, DPO) and parameter-agnostic techniques (e.g., prompt engineering, retrieval-augmented generation, reward shaping, trajectory construction). It further analyzes critical integrative aspects—such as hybrid optimization—and synthesizes evaluation benchmarks and representative applications. Contributions include: (1) a structured, comprehensive review encompassing over 100 works; (2) an open-source, standardized reference library hosted on GitHub; and (3) a clear articulation of open challenges and actionable research directions. Collectively, this work provides both theoretical foundations and a reproducible toolchain for efficient LLM agent optimization.
To address the weak planning capability and poor generalization of LLM-based agents, this paper proposes AgentGen—a novel framework for agent instruction tuning. First, it constructs a domain-inspired heuristic environment model that automatically generates structurally diverse simulated environments from an inspiration corpus. Second, it introduces a Bidirectional Evolution (Bi-Evol) algorithm that jointly optimizes task difficulty progression—forward-growing task complexity and backward validation—to produce smooth, pedagogically grounded task sequences. Third, it performs environment-task co-driven instruction fine-tuning. Experiments show that AgentGen-finetuned Llama-3.1-8B outperforms GPT-3.5, while Llama-3.1-70B achieves state-of-the-art performance across multi-domain planning benchmarks, with significant gains in stepwise reasoning and cross-environment generalization. The core contributions are: (i) a principled mechanism for generating diverse, semantically rich environments, and (ii) a bidirectional task evolution paradigm that bridges curriculum learning and robust generalization.
Current LLM-based algorithm design relies heavily on empirical trial-and-error, lacking formal theoretical foundations to systematically analyze how critical design choices—such as task decomposition strategies and prompt engineering—affect accuracy and computational efficiency. Method: We propose the first formal analytical framework for LLM-invocation algorithms, modeling LLM subroutines as a computational graph, establishing structured principles for task decomposition, and introducing an error propagation model that enables provable analysis of both accuracy and computational complexity. Contribution/Results: Empirically validated across parallel, hierarchical, and recursive paradigms, our framework explains observed empirical phenomena, guides prompt design and granularity selection, predicts performance bottlenecks, and inspires novel robust algorithm designs. The implementation is publicly available.
This work addresses the limitations of large language models (LLMs) in optimizing actions over long-horizon sequential decision-making tasks and the inability of reinforcement learning (RL) to perform high-level abstraction and task decomposition. To bridge this gap, the paper proposes a hierarchical hybrid architecture that, for the first time, effectively integrates the semantic reasoning capabilities of LLMs with the precise control mechanisms of RL. In this framework, the LLM generates subgoals and structured plans, while the RL agent optimizes low-level action policies through environmental interaction. Evaluated across multiple sequential decision tasks, the approach significantly improves sample efficiency, task success rates, and trajectory coherence, outperforming both pure RL and pure LLM baselines, thereby demonstrating the efficacy of synergistically combining high-level planning with low-level execution.
This work addresses the limitations of existing large language model (LLM) agents, which typically employ fixed-granularity planning mechanisms that struggle to balance efficiency on simple tasks with the detailed reasoning required for complex ones. To overcome this, the paper introduces AdaPlan-H, a cognitively inspired adaptive hierarchical planning framework that, for the first time, integrates a progressive refinement strategy into LLM agents to enable dynamic, task-difficulty-aware adjustment of planning granularity. By synergistically combining hierarchical task decomposition, imitation learning, and capability enhancement, AdaPlan-H supports the adaptive generation and continuous optimization of planning hierarchies. Experimental results demonstrate that AdaPlan-H significantly improves success rates on multi-step complex tasks while effectively avoiding over-planning, thereby validating its efficiency and flexibility.
This work addresses the challenge of reliably decomposing ambiguous or temporally extended natural language instructions into executable actions for heterogeneous robots in multi-robot task planning, where existing approaches often lack robustness and feasibility. The authors propose a hierarchical multi-agent LLM-based planning framework: a high-level module performs task decomposition and assignment, while low-level agents generate PDDL problem instances solved by classical planners. Upon planning failure, the system employs TextGrad-inspired textual gradient optimization to refine prompts and shares meta-prompts among peer agents to enhance efficiency. Evaluated on the MAT-THOR benchmark, the method achieves success rates of 0.95, 0.84, and 0.60 on composite, complex, and ambiguous tasks, respectively—outperforming state-of-the-art baselines by 2–15 percentage points. Ablation studies confirm the contribution of each component to overall performance.
This study investigates whether large language models (LLMs) can effectively replace or collaborate with classical symbolic planners for task planning. To this end, the authors propose a novel paradigm that integrates a PDDL-based planner with an LLM via the Model Context Protocol (MCP): actions from PyPDDLEngine are exposed as callable tools, enabling the LLM to generate plans interactively and incrementally based on environmental feedback, in contrast to conventional one-shot plan generation. Evaluated on 102 Blocksworld tasks, the LLM agent achieved a planning success rate of 66.7%, slightly outperforming direct prompting (63.7%) but at 5.7 times higher token cost. Both approaches produced notably shorter plans, suggesting reliance on memorized patterns from training data rather than robust generalization. This work offers a fine-grained interaction framework for synergizing LLMs with symbolic planning.
Long-horizon hierarchical planning in text-based environments faces challenges including open-ended action spaces, ambiguous observations, and sparse rewards; existing LLM-dependent approaches suffer from high inference overhead, non-differentiable parameters, and inefficient deployment. Method: We propose the “One-Shot Teacher” paradigm: an LLM is invoked only once at planning initialization to generate a subgoal sequence, followed by LLM-guided trajectory distillation to pretrain a lightweight student planner (e.g., a Transformer-based planner) for subgoal-conditioned modeling. Contribution/Results: This eliminates repeated LLM calls during both training and inference. On TextCraft, our method achieves 56% success rate—surpassing ADaPT’s 52%—while reducing average inference time from 164.4 seconds to 3.0 seconds (54× speedup), significantly improving both task performance and deployment efficiency.