llm-guided action synthesis

Designs and implements systems that use large language models to interpret contextual or multimodal prompts and translate semantics into ordered sequences of action primitives; these systems produce executable plans and generate control or dispatch commands to execution queues or interfaces. Competence includes language-model–level planning, mapping intents to low-level actions, sequencing those actions into executable plans, and interfacing plan outputs with runtime dispatch mechanisms.

llm-guidedactionsynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

A Survey on Large Language Models for Automated Planning

Feb 18, 2025
MA
Mohamed Aghzal
🏛️ George Mason University | National Science Foundation

This work investigates the fundamental applicability boundaries of large language models (LLMs) in automated planning. Through a systematic literature review, a multi-dimensional capability assessment framework, and empirical evaluation on canonical domains—including Block World and Logistics—the study reveals critical limitations: inconsistent long-horizon reasoning, failure in constraint-sensitive planning, and unreliable state tracking. Methodologically, it employs rigorous comparative analysis across diverse planning tasks to isolate intrinsic LLM deficiencies. The primary contribution is the first principled argument that LLMs are unsuitable as standalone planners; instead, it proposes “hybrid intelligent planning”—a novel paradigm wherein LLMs serve exclusively as semantic understanding and heuristic generation modules, tightly integrated with symbolic reasoning engines and search algorithms. The work establishes a reproducible, taxonomy-based evaluation methodology and provides concrete architectural design principles for synergistic LLM–symbolic system integration, thereby delivering both theoretical foundations and practical guidelines for LLM-augmented planning.

Assess LLMs' planning capabilitiesIdentify limitations in long-horizon reasoningPropose hybrid LLM-traditional planning methods

Large Language Models for Planning: A Comprehensive and Systematic Survey

May 26, 2025
PC
Pengfei Cao
🏛️ University of Chinese Academy of Sciences | Harbin Institute of Technology | Institute of Information Engineering | Chinese Academy of Sciences | Beijing Institute of Technology

Despite growing interest in leveraging large language models (LLMs) for planning—requiring environmental understanding, logical reasoning, and sequential decision-making—there exists no systematic taxonomy or standardized evaluation framework. Method: This paper introduces the first unified classification scheme for LLM-based planning methods, categorizing existing approaches into three paradigms: external module augmentation, fine-tuning-driven methods, and search-oriented techniques. It further establishes a standardized evaluation framework encompassing benchmark tasks, multidimensional metrics, and empirical comparisons. Contribution/Results: Through comprehensive literature analysis, methodological abstraction, and cross-paradigm mechanistic synthesis, this work delivers the field’s first holistic survey. It clarifies the technical evolution trajectory, identifies core bottlenecks—including scalability, generalization, and causal reasoning—and proposes future directions such as trustworthy planning, embodied collaboration, and neuro-symbolic integration. The study provides an authoritative knowledge graph and methodological roadmap for advancing LLM-based planning research.

Investigating LLMs' broader application in intelligent agent planningReviewing three principal LLM-based planning methodologiesSummarizing evaluation frameworks for LLM-based planning performance

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of translating natural language intents into robot-executable actions within dynamic, unknown environments by proposing an intent-driven dual-AI collaborative framework. The framework leverages large language models to generate constrained executable code and integrates vision-language models for semantic grounding. Its core innovation lies in an adaptive replanning mechanism triggered by geometric and semantic thresholds, which achieves closed-loop control through runtime monitoring. Experimental results demonstrate that the proposed approach robustly executes complex instructions under bounded reaction cycles, enabling effective robot control in dynamic settings. These findings validate the reliability and generalization capability of generative AI for real-time embodied intelligence tasks.

Adaptive Code GenerationDynamic EnvironmentsIntention-based Autonomy

Plan-and-Act using Large Language Models for Interactive Agreement

Apr 01, 2025
KS
Kazuhiro Sasabuchi
🏛️ Microsoft

This work addresses the fundamental tension in LLM-driven robotics for dynamic human–robot interaction (HRI): balancing respect for the human’s ongoing activity with efficient task execution. We propose a temporally aware “plan–execute” skill framework featuring a novel two-stage LLM invocation mechanism: (1) an initial LLM call generates a high-level action plan; (2) a second, context-triggered invocation is autonomously scheduled based on real-time HRI state—specifically, the robot’s current action description—enabling adaptive switching between passive responsiveness and proactive intervention. The framework integrates prompt engineering, temporal reasoning, explicit HRI state modeling, and Engage skill composition. Evaluated across four heterogeneous real-world HRI scenarios, our approach achieves a 90% task success rate, demonstrating substantial improvements in behavioral appropriateness, timing accuracy, and cross-scenario generalizability of LLM-driven robotic agents.

Balancing human activity respect and robot task priority in HRIDetermining optimal timing for LLM-based action planningScaling LLM application across diverse human-robot interaction scenarios

This study addresses security vulnerabilities in LLM-driven embodied agents by conceptualizing environmental states as an attack surface and introducing the novel concept of "state semantic injection." Through the development of a comprehensive framework encompassing state semantic modeling, adversarial injection testing, and security evaluation, this work systematically elucidates the mechanisms by which malicious state exploitation induces task execution deviations. Experimental results validate that such attacks can trigger behavioral anomalies, thereby establishing a new class of security risks. Consequently, this research not only expands the threat model for embodied intelligence but also provides a critical theoretical foundation and fresh perspectives for enhancing system robustness and developing effective defense mechanisms against emerging adversarial threats.

Attack SurfaceEmbodied AgentsLarge Language Models

Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks

Mar 12, 2025
LE
Lutfi Eren Erdogan
🏛️ UC Berkeley | University of Tokyo | ICSI

To address the weak planning capability and low plan accuracy of large language models (LLMs) in long-horizon, multi-step tasks, this paper proposes a planner-executor decoupled two-stage framework: a Planner generates structured high-level plans, while an Executor performs environment-specific actions. We introduce an explicit planning augmentation paradigm, designing a scalable synthetic data generation method to construct diverse plan trajectories with ground-truth annotations. By integrating trajectory alignment annotation, synthetic data distillation, and generalization-enhanced training, we significantly improve planning robustness. Evaluated on the WebArena-Lite benchmark, our approach achieves a 54% task success rate—setting a new state-of-the-art for long-horizon web navigation—and establishes a novel paradigm for reliable long-term planning in LLM-based agents.

Achieving high success rates in web navigation tasksEnhancing plan generation with synthetic dataImproving planning for long-horizon tasks using LLMs

Bootstrapping Object-level Planning with Large Language Models

Sep 18, 2024
DP
D. Paulius
🏛️ Brown University | University of Innsbruck

Direct generation of PDDL goals or task sequences by large language models (LLMs) often yields semantically abstract, non-executable outputs, hindering integration with robot task and motion planning (TAMP). Method: We propose an LLM knowledge distillation framework that extracts object-level state-change knowledge via prompt engineering, constructs a function-oriented object network (FOON), and automatically compiles it into semantically aligned, executable PDDL subgoals. Contribution/Results: This FOON-PDDL joint representation establishes the first structured synergy between LLM-derived high-level semantics and classical planners’ action-object constraints. Evaluated on simulated pick-and-place tasks, our approach improves subgoal success rate by 37%, significantly enhancing planning feasibility and cross-task generalization.

Extracts knowledge from LLMs for object-level planning.Generates PDDL subgoals from functional object-oriented networks.Improves task and motion planning in pick-and-place tasks.

Latest Papers

What's happening recently
View more

This work addresses the challenge that existing vision-language-action models struggle to accurately execute natural language instructions involving spatiotemporal and logical constraints, while also lacking interpretability. The authors propose a hierarchical framework that, for the first time, deeply integrates Signal Temporal Logic (STL) between language understanding and robotic execution. The approach decomposes high-level instructions into subtasks and generates verifiable, optimizable, and correctable STL specifications, which dynamically schedule low-level policies. By combining vision-language models, STL, model predictive control, and learned policies, the method enables an end-to-end mapping from natural language instructions to formal specifications, supporting online monitoring and replanning. Experiments in real-world tabletop environments demonstrate significant improvements in accuracy, reliability, and interpretability of language-guided robotic tasks.

interpretabilitynatural language instructionsprecise specification

Current large language model (LLM) agents predominantly rely on informal prompts to encode skills, lacking built-in support for workflow state, execution policies, and completion criteria, which often leads to a disconnect between reasoning and action. This work proposes Formal Skill—a runtime-native, reusable capability abstraction that encodes skills as executable state machines. Each skill is defined via a JSON Schema interface, implemented with a Python executor for action logic, and enhanced with event-driven hooks and local state management. This approach represents the first systematic shift from prompt-based skill descriptions to stateful, policy-constrained executable units. By doing so, it substantially improves skill composability, observability, and policy enforcement. Evaluated on Harness-Bench, the method achieves highly competitive performance with significantly fewer token expenditures, particularly excelling in tasks requiring structured skill dependencies.

formal skillsLLM agentsruntime abstraction

This work addresses the limited capability of service and assistant robots in task and motion planning during natural language interaction by proposing a hierarchical language-driven framework that decouples high-level task planning from low-level spatial reasoning through the collaboration of two large language models (LLMs). The high-level agent interprets natural language instructions to generate action sequences, while the low-level module integrates YOLOX-GDRNet for object detection and pose estimation, employing ReAct-style prompting and tool-calling mechanisms to handle 3D spatial placement tasks and identify infeasible requests. Evaluated across 24 test scenarios ranging from simple to complex instructions, the system achieves an end-to-end task success rate of 86%, significantly enhancing the intuitiveness and robustness of human-robot collaboration.

hierarchical promptinghuman-robot interactionnatural language commands

Hot Scholars

YL

Yang Liu

Dalian University of Technology
computer visionimage processing
VN

Vu Nguyen

Stony Brook University
Computer VisionMachine Learning
CS

Chenxi Song

Westlake University & Jilin University
3DVIsion3D&4D Generation&Reconstruction
BJ

Baoxiong Jia

Ph.D. in Computer Science, UCLA
Computer VisionArtificial Intelligence
CK

Claudius Kienle

Intelligent Autonomous Systems / TU Darmstadt
Machine LearningRobotics