Score
Designs and implements systems that use large language models to interpret contextual or multimodal prompts and translate semantics into ordered sequences of action primitives; these systems produce executable plans and generate control or dispatch commands to execution queues or interfaces. Competence includes language-model–level planning, mapping intents to low-level actions, sequencing those actions into executable plans, and interfacing plan outputs with runtime dispatch mechanisms.
Existing research on large language models (LLMs) as autonomous agents and tool users remains fragmented and limited in architecture design, multi-agent coordination, tool integration, cognitive mechanism modeling, and evaluation frameworks. Method: This survey systematically analyzes 2023–2025 top-tier conference and journal publications using structured literature analysis, integrating prompt engineering and fine-tuning techniques to dissect LLM implementations of core cognitive capabilities—reasoning, planning, and memory. Contribution/Results: We identify three breakthrough directions—verifiable reasoning, self-improvement, and personalized customization—and distill ten concrete future research pathways. Further, we propose a unified evaluation framework covering 68 publicly available datasets, exposing critical gaps in current benchmarks regarding task generalization, dynamic adaptability, and causal attribution capability.
This work investigates the fundamental applicability boundaries of large language models (LLMs) in automated planning. Through a systematic literature review, a multi-dimensional capability assessment framework, and empirical evaluation on canonical domains—including Block World and Logistics—the study reveals critical limitations: inconsistent long-horizon reasoning, failure in constraint-sensitive planning, and unreliable state tracking. Methodologically, it employs rigorous comparative analysis across diverse planning tasks to isolate intrinsic LLM deficiencies. The primary contribution is the first principled argument that LLMs are unsuitable as standalone planners; instead, it proposes “hybrid intelligent planning”—a novel paradigm wherein LLMs serve exclusively as semantic understanding and heuristic generation modules, tightly integrated with symbolic reasoning engines and search algorithms. The work establishes a reproducible, taxonomy-based evaluation methodology and provides concrete architectural design principles for synergistic LLM–symbolic system integration, thereby delivering both theoretical foundations and practical guidelines for LLM-augmented planning.
Despite growing interest in leveraging large language models (LLMs) for planning—requiring environmental understanding, logical reasoning, and sequential decision-making—there exists no systematic taxonomy or standardized evaluation framework. Method: This paper introduces the first unified classification scheme for LLM-based planning methods, categorizing existing approaches into three paradigms: external module augmentation, fine-tuning-driven methods, and search-oriented techniques. It further establishes a standardized evaluation framework encompassing benchmark tasks, multidimensional metrics, and empirical comparisons. Contribution/Results: Through comprehensive literature analysis, methodological abstraction, and cross-paradigm mechanistic synthesis, this work delivers the field’s first holistic survey. It clarifies the technical evolution trajectory, identifies core bottlenecks—including scalability, generalization, and causal reasoning—and proposes future directions such as trustworthy planning, embodied collaboration, and neuro-symbolic integration. The study provides an authoritative knowledge graph and methodological roadmap for advancing LLM-based planning research.
This study addresses the challenge of translating natural language intents into robot-executable actions within dynamic, unknown environments by proposing an intent-driven dual-AI collaborative framework. The framework leverages large language models to generate constrained executable code and integrates vision-language models for semantic grounding. Its core innovation lies in an adaptive replanning mechanism triggered by geometric and semantic thresholds, which achieves closed-loop control through runtime monitoring. Experimental results demonstrate that the proposed approach robustly executes complex instructions under bounded reaction cycles, enabling effective robot control in dynamic settings. These findings validate the reliability and generalization capability of generative AI for real-time embodied intelligence tasks.
This work addresses the fundamental tension in LLM-driven robotics for dynamic human–robot interaction (HRI): balancing respect for the human’s ongoing activity with efficient task execution. We propose a temporally aware “plan–execute” skill framework featuring a novel two-stage LLM invocation mechanism: (1) an initial LLM call generates a high-level action plan; (2) a second, context-triggered invocation is autonomously scheduled based on real-time HRI state—specifically, the robot’s current action description—enabling adaptive switching between passive responsiveness and proactive intervention. The framework integrates prompt engineering, temporal reasoning, explicit HRI state modeling, and Engage skill composition. Evaluated across four heterogeneous real-world HRI scenarios, our approach achieves a 90% task success rate, demonstrating substantial improvements in behavioral appropriateness, timing accuracy, and cross-scenario generalizability of LLM-driven robotic agents.
This study addresses security vulnerabilities in LLM-driven embodied agents by conceptualizing environmental states as an attack surface and introducing the novel concept of "state semantic injection." Through the development of a comprehensive framework encompassing state semantic modeling, adversarial injection testing, and security evaluation, this work systematically elucidates the mechanisms by which malicious state exploitation induces task execution deviations. Experimental results validate that such attacks can trigger behavioral anomalies, thereby establishing a new class of security risks. Consequently, this research not only expands the threat model for embodied intelligence but also provides a critical theoretical foundation and fresh perspectives for enhancing system robustness and developing effective defense mechanisms against emerging adversarial threats.
To address the weak planning capability and low plan accuracy of large language models (LLMs) in long-horizon, multi-step tasks, this paper proposes a planner-executor decoupled two-stage framework: a Planner generates structured high-level plans, while an Executor performs environment-specific actions. We introduce an explicit planning augmentation paradigm, designing a scalable synthetic data generation method to construct diverse plan trajectories with ground-truth annotations. By integrating trajectory alignment annotation, synthetic data distillation, and generalization-enhanced training, we significantly improve planning robustness. Evaluated on the WebArena-Lite benchmark, our approach achieves a 54% task success rate—setting a new state-of-the-art for long-horizon web navigation—and establishes a novel paradigm for reliable long-term planning in LLM-based agents.
Direct generation of PDDL goals or task sequences by large language models (LLMs) often yields semantically abstract, non-executable outputs, hindering integration with robot task and motion planning (TAMP). Method: We propose an LLM knowledge distillation framework that extracts object-level state-change knowledge via prompt engineering, constructs a function-oriented object network (FOON), and automatically compiles it into semantically aligned, executable PDDL subgoals. Contribution/Results: This FOON-PDDL joint representation establishes the first structured synergy between LLM-derived high-level semantics and classical planners’ action-object constraints. Evaluated on simulated pick-and-place tasks, our approach improves subgoal success rate by 37%, significantly enhancing planning feasibility and cross-task generalization.
This work addresses the challenge that existing vision-language-action models struggle to accurately execute natural language instructions involving spatiotemporal and logical constraints, while also lacking interpretability. The authors propose a hierarchical framework that, for the first time, deeply integrates Signal Temporal Logic (STL) between language understanding and robotic execution. The approach decomposes high-level instructions into subtasks and generates verifiable, optimizable, and correctable STL specifications, which dynamically schedule low-level policies. By combining vision-language models, STL, model predictive control, and learned policies, the method enables an end-to-end mapping from natural language instructions to formal specifications, supporting online monitoring and replanning. Experiments in real-world tabletop environments demonstrate significant improvements in accuracy, reliability, and interpretability of language-guided robotic tasks.
研究探讨了通过激活空间中的向量表示来操纵大型语言模型的程序技能,发现这些技能可以作为方向被激活和组合,以实现更高级别的功能和个人化优化。
Current large language model (LLM) agents predominantly rely on informal prompts to encode skills, lacking built-in support for workflow state, execution policies, and completion criteria, which often leads to a disconnect between reasoning and action. This work proposes Formal Skill—a runtime-native, reusable capability abstraction that encodes skills as executable state machines. Each skill is defined via a JSON Schema interface, implemented with a Python executor for action logic, and enhanced with event-driven hooks and local state management. This approach represents the first systematic shift from prompt-based skill descriptions to stateful, policy-constrained executable units. By doing so, it substantially improves skill composability, observability, and policy enforcement. Evaluated on Harness-Bench, the method achieves highly competitive performance with significantly fewer token expenditures, particularly excelling in tasks requiring structured skill dependencies.
研究通过约束大语言模型和物理验证,解决机器人在复杂环境中安全执行任务的问题,提高操作成功率。
This work addresses the limited capability of service and assistant robots in task and motion planning during natural language interaction by proposing a hierarchical language-driven framework that decouples high-level task planning from low-level spatial reasoning through the collaboration of two large language models (LLMs). The high-level agent interprets natural language instructions to generate action sequences, while the low-level module integrates YOLOX-GDRNet for object detection and pose estimation, employing ReAct-style prompting and tool-calling mechanisms to handle 3D spatial placement tasks and identify infeasible requests. Evaluated across 24 test scenarios ranging from simple to complex instructions, the system achieves an end-to-end task success rate of 86%, significantly enhancing the intuitiveness and robustness of human-robot collaboration.