transferable strategy learning

Designs and implements models and representations that capture high-level planning or decision-making strategies separately from low-level action execution, along with decoders and tuning procedures to map those representations to concrete executors. Builds training and evaluation methods to learn reusable strategy representations, transfer them across tasks and agents, and integrate strategy-grounded planners or planners trained independently of executors to produce appropriate action sequences.

transferablestrategylearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Large Language Models for Planning: A Comprehensive and Systematic Survey

May 26, 2025
PC
Pengfei Cao
🏛️ University of Chinese Academy of Sciences | Harbin Institute of Technology | Institute of Information Engineering | Chinese Academy of Sciences | Beijing Institute of Technology

Despite growing interest in leveraging large language models (LLMs) for planning—requiring environmental understanding, logical reasoning, and sequential decision-making—there exists no systematic taxonomy or standardized evaluation framework. Method: This paper introduces the first unified classification scheme for LLM-based planning methods, categorizing existing approaches into three paradigms: external module augmentation, fine-tuning-driven methods, and search-oriented techniques. It further establishes a standardized evaluation framework encompassing benchmark tasks, multidimensional metrics, and empirical comparisons. Contribution/Results: Through comprehensive literature analysis, methodological abstraction, and cross-paradigm mechanistic synthesis, this work delivers the field’s first holistic survey. It clarifies the technical evolution trajectory, identifies core bottlenecks—including scalability, generalization, and causal reasoning—and proposes future directions such as trustworthy planning, embodied collaboration, and neuro-symbolic integration. The study provides an authoritative knowledge graph and methodological roadmap for advancing LLM-based planning research.

Investigating LLMs' broader application in intelligent agent planningReviewing three principal LLM-based planning methodologiesSummarizing evaluation frameworks for LLM-based planning performance

Must-Read Papers

Most classic and influential ideas
View more

A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks

Oct 07, 2025
SS
Shuzheng Si
🏛️ Tsinghua University | University of Illinois Urbana-Champaign | DeepLang AI | Peking University

To address the blind trial-and-error behavior and hallucinated actions exhibited by large language model (LLM) agents in long-horizon tasks—stemming from insufficient global planning capability—this paper proposes EAGLET. Our method operates in two stages: first, it automatically generates high-quality planning data via a homologous consensus filtering mechanism; second, it employs execution-ability-gain-driven, rule-guided reinforcement learning—requiring no human annotation. EAGLET integrates state-of-the-art LLM-based planning generation, supervised fine-tuning for cold-start initialization, and interpretable reward modeling to construct a plug-and-play, lightweight global planner. Evaluated on three long-horizon benchmarks, EAGLET achieves state-of-the-art performance while reducing training cost by 8× compared to standard RL baselines. It significantly enhances both planning reliability and training efficiency.

Enhancing global planning for long-horizon agent tasksReducing hallucinatory actions and trial-and-error in LLM agentsTraining efficient planners without human effort or extra data

Current LLM-driven web agents often suffer from inadequate planning strategies, leading to poor exploration, omission of critical steps, and high sensitivity to task constraints. This work proposes PlanAhead, a framework that systematically evaluates—for the first time—the impact of four natural language planning representations (subgoal sequences, narrative descriptions, pseudocode, and checklists) on the performance of multimodal LLM agents. It introduces an innovative, annotation-free three-tier automatic task difficulty classification method and proposes two novel evaluation metrics: Achievement Rate (AR) and Solved-Task Consistency (STC). Experimental results demonstrate that both the choice of planning representation and the underlying LLM significantly affect agent robustness and success rates, with pronounced differences especially evident in high-difficulty tasks.

agent robustnessLLM web agentsplan formulation

Bootstrapping Object-level Planning with Large Language Models

Sep 18, 2024
DP
D. Paulius
🏛️ Brown University | University of Innsbruck

Direct generation of PDDL goals or task sequences by large language models (LLMs) often yields semantically abstract, non-executable outputs, hindering integration with robot task and motion planning (TAMP). Method: We propose an LLM knowledge distillation framework that extracts object-level state-change knowledge via prompt engineering, constructs a function-oriented object network (FOON), and automatically compiles it into semantically aligned, executable PDDL subgoals. Contribution/Results: This FOON-PDDL joint representation establishes the first structured synergy between LLM-derived high-level semantics and classical planners’ action-object constraints. Evaluated on simulated pick-and-place tasks, our approach improves subgoal success rate by 37%, significantly enhancing planning feasibility and cross-task generalization.

Extracts knowledge from LLMs for object-level planning.Generates PDDL subgoals from functional object-oriented networks.Improves task and motion planning in pick-and-place tasks.

TwoStep: Multi-agent Task Planning using Classical Planners and Large Language Models

Mar 25, 2024
IS
Ishika Singh
🏛️ University of Southern California

To address the challenge of balancing concurrent action conflicts and plan executability in multi-agent task planning, this paper proposes a two-stage LLM-PDDL collaborative framework. First, a large language model (LLM) performs commonsense-driven goal decomposition to generate mutually exclusive, parallelizable sub-goals. Second, a classical PDDL planner (e.g., FF or Fast Downward) independently synthesizes formally verifiable single-agent plans for each agent. This work is the first to integrate the LLM’s high-level goal abstraction capability with the formal correctness guarantees of classical planning. Empirical results demonstrate 100% action executability, significantly reduced planning time, and plan step counts that outperform single-agent baselines while approaching human expert performance—thereby unifying efficiency, feasibility, and coordination quality in multi-agent planning.

Combining classical planning and LLMs for multi-agent task decompositionImproving planning time and execution steps in multi-agent PDDLMatching human expert subgoal performance with LLM approximations

Improving Planning with Large Language Models: A Modular Agentic Architecture

Sep 30, 2023
TW
Taylor Webb
🏛️ Microsoft Research | Princeton University

Large language models (LLMs) exhibit limited planning capabilities in multi-step reasoning and goal-directed tasks. To address this, we propose the Modular Agent Planning (MAP) architecture—a cognitively inspired, reinforcement learning–informed framework that decomposes planning into specialized LLM modules: conflict monitoring, state prediction, and task decomposition. These modules operate in a recurrent, collaborative loop to enable dynamic, adaptive planning. MAP supports lightweight deployment and cross-task generalization without fine-tuning, seamlessly adapting to diverse LLM scales (e.g., Llama3-70B). Empirical evaluation on graph traversal, Tower of Hanoi, PlanBench, and StrategyQA demonstrates substantial improvements over zero-shot prompting, chain-of-thought, and tree-of-thought baselines—yielding higher planning accuracy and robustness. Our core contribution is the first systematic integration of modular, division-of-labor mechanisms into LLM-based planning, establishing a novel paradigm for structured, interpretable, and scalable agent reasoning.

Addresses multi-step reasoning limitations in large language modelsEnhances planning via specialized modules interacting through LLM callsProposes modular architecture for improved goal-directed planning tasks

Latest Papers

What's happening recently
View more

Existing language model agents struggle to efficiently execute complex instructions in long-horizon tasks due to insufficient planning capabilities. This work proposes a planner-centric multi-agent framework comprising a planner, an executor, and a memory manager. Through computational resource allocation analysis, we demonstrate that the planning component predominantly governs overall performance. Leveraging this insight, we apply reinforcement learning exclusively to the planner, incorporating trajectory-level rewards and a vision-language model-based evaluation mechanism to enable asymmetric computation allocation. The resulting approach achieves significant performance gains across diverse benchmarks—including web navigation, operating system control, and tool usage—thereby validating the efficacy and strong generalization of prioritizing high-level planning.

language model agentslong-horizon planningmulti-agent collaboration

This work addresses performance bottlenecks in multi-agent reinforcement learning (MARL) arising from sparse rewards, high-dimensional state-action spaces, and the challenge of coordinated policy learning. The authors propose a hierarchical architecture wherein a pretrained large language model (LLM) serves as a centralized strategic controller at the high level, dynamically selecting among specialized low-level reinforcement learning policies without relying on handcrafted rules. This approach represents the first integration of LLMs into high-level planning for multi-agent systems, significantly enhancing tactical diversity and behavioral human-likeness. Evaluated on a 2v2 capture-the-flag task, the method achieves a win rate of 46.4%, matching the performance of hand-designed behavior trees and substantially outperforming flat RL baselines. A user study further reveals that 60% of participants judged the agents’ behavior as most human-like (p = 0.027).

coordinated strategieslarge state-action spacesmulti-agent reinforcement learning

This work addresses the vulnerability of external skill updating to sparse or noisy execution trajectories, which can lead to the entrenchment of ineffective or even detrimental policies. To mitigate this issue, the authors propose a training-free skill optimization framework that leverages a falsifiable hypothesis-driven mechanism. By integrating controlled experiments, behavioral discrepancy analysis, and progressive skill disclosure, the method enables auditable and noise-resilient skill curation and execution at frozen model inference endpoints. Evaluated on ALFWorld, the approach yields substantial performance gains: average success rates improve by 6.9 and 4.0 percentage points for Qwen3-8B and Qwen3.6-27B, respectively. Notably, it maintains a +7.1-point advantage even under 20% erroneous feedback and demonstrates strong cross-run and cross-model transferability.

frozen modelshypothesis-drivenLLM agents

This work investigates how intelligent agents can dynamically balance fast reactive control against slower yet more robust deliberative planning to achieve both efficiency and performance. To this end, the authors propose a learnable meta-reasoning controller trained via reinforcement learning that adaptively triggers planning based on an uncertainty score derived from the reactive policy. The framework integrates reinforcement learning, imitation learning, model-based planning, and uncertainty estimation, enabling the agent to progressively shift toward purely reactive control as the reactive policy improves. Experiments in motion planning and navigation tasks demonstrate that the agent accurately discerns when to rely on reactive responses versus when to invoke planning, dynamically optimizing computational resource allocation throughout training and yielding a flexible, efficient decision-making architecture.

computation allocationdecision-makingdeliberative planning

Hot Scholars

TA

The Anh Han

Professor of Computer Science, Teesside University
Evolutionary Game TheoryArtificial IntelligenceEvolution of CooperationMulti-agent Systems
SZ

Shufang Zhu

University of Liverpool
Artificial IntelligenceFormal MethodsReactive Synthesis
MM

Munyque Mittelmann

CNRS, LIPN, Université Sorbonne Paris Nord
Multi-Agent SystemsFormal MethodsStrategic ReasoningModal Logic
GF

Gabriele Farina

Assistant Professor of Computer Science, MIT
Computational Game TheoryOptimizationEconomics and Computation