Score
Designs and builds planners that translate natural-language or multimodal instructions into executable action sequences or compact latent plan representations encoding semantic intent and goal conditions. These systems ground language to perceived objects, spatial relations, and targets, produce instruction rewrites or grounding cues, and incorporate constraints (e.g., accessibility, social norms) using techniques such as latent reasoning tokens and mutual-information objectives for goal-conditioned semantic planning.
This work addresses a critical gap in embodied intelligence research, where the purported role of language often lacks empirical grounding, making it difficult to assess its actual contribution to agent behavior. To bridge this disconnect, the paper introduces a functional-role-based analytical framework that categorizes language’s roles in embodied systems into five distinct types: Specification, Embodied Representation, Action Orchestration, Grounding Regulation, and Execution Coupling. This framework enables a systematic evidence audit of existing approaches and is the first to uniformly apply across both modular and end-to-end architectures, facilitating fine-grained evaluation. The analysis reveals that most current studies fail to effectively leverage language as a mediating mechanism or provide sufficient evidence linking linguistic components to behavioral improvements, with performance gains frequently lacking rigorous attribution to language use.
This work addresses the challenges of adaptability and robustness in robotic execution of natural language instructions within dynamic environments. We propose a deep synergy framework integrating large language models (LLMs) with semantic digital twins (SDTs). Methodologically, we introduce the first bidirectional coupling: the SDT provides real-time, semantically grounded environmental representations and affordance-aware perception, while the LLM performs instruction parsing, reflective reasoning, and generation of structured action triplets—augmented with failure-driven iterative replanning. Our technical contributions include SDT modeling, adaptation to the ALFRED benchmark, and closed-loop simulation evaluation. Experiments demonstrate significant improvements in task success rate and fault tolerance across diverse household scenarios, effectively handling uncertainties such as object pose variations, occlusions, and execution failures. The framework establishes a novel paradigm for semantic-level task planning in embodied intelligence.
Diffusion Transformers often suffer from visual hallucinations and poor instruction alignment in complex video generation—particularly for high-level semantic tasks involving human-object interaction, multi-stage actions, and contextual motion reasoning. To address these limitations, we propose Plan-X, the first framework to introduce a learnable multimodal semantic planning mechanism. Plan-X employs a multimodal large language model as a semantic planner that jointly reasons over textual and visual context to infer user intent and autoregressively generate spatiotemporal semantic tokens. These tokens serve as structured priors that explicitly guide the diffusion Transformer toward high-fidelity video synthesis. Experiments demonstrate that Plan-X significantly suppresses visual hallucinations, improves instruction adherence and spatiotemporal consistency, and achieves state-of-the-art performance on complex scene understanding and long-horizon action modeling tasks.
Current large model–based approaches to autonomous driving scene understanding and planning lack effective temporal modeling, leading to inconsistent reasoning over sequential actions and compromising both safety and interpretability. To address this, this work proposes three multi-agent planner architectures incorporating varying degrees of temporal conditioning constraints. The authors establish the first empirical benchmark for temporally aware scene-to-planning reasoning on a subset of BDD-X and introduce evaluation metrics assessing semantic, syntactic, and logical consistency. Experimental results show that while explicit temporal constraints do not significantly improve standard NLP metrics, qualitative analysis reveals their capacity to elicit forward-looking risk assessment, stabilize corrective behaviors, and enhance strategic diversity. The study also highlights limitations in current prompt engineering practices regarding temporal grounding.
To address the challenges of zero-shot planning difficulty and weak dynamic adaptability in Embodied Instruction Following (EIF), this paper proposes the first training-free, vision-driven zero-shot embodied planning framework. Methodologically, it decomposes natural language instructions into executable high-level sub-goal sequences via a self-questioning-and-answering mechanism; introduces a visually grounded real-time re-planning mechanism that dynamically refines plans based on visual feedback during interaction; and designs RelaxedHLP—a novel evaluation metric that quantifies high-level planning quality for the first time. Experiments on the ALFRED benchmark demonstrate state-of-the-art zero-shot and few-shot performance, particularly on complex tasks requiring multi-step reasoning and environment responsiveness. Visual feedback significantly improves re-planning accuracy, with consistent gains over existing methods.
Existing robotic navigation systems struggle with unstructured, map-free environments when receiving incomplete natural-language task instructions. Method: This work proposes an online semantic planning framework leveraging large language models (LLMs), integrated into a closed-loop system that synergistically combines real-time semantic SLAM, receding-horizon planning (RHP), and online safety verification. The framework enables concurrent semantic mapping, automatic task completion, dynamic subtask re-planning, and runtime safety constraint enforcement—without requiring prior maps or manually refined commands. Contribution/Results: It establishes the first end-to-end online pipeline for semantic understanding, planning, and execution. Evaluated in cluttered outdoor environments exceeding 20,000 m², the approach reduces task completion time and path length by over 50% compared to baselines, while significantly decreasing user interaction frequency.
Robots frequently fail to execute everyday tasks due to natural language instructions lacking commonsense preconditions and decomposed subgoals. This paper proposes an LLM-augmented symbolic planning framework that, for the first time, automatically formalizes implicit preconditions and refined subgoals generated by large language models (LLMs) and integrates them into classical planners (e.g., PDDL), enabling end-to-end instruction-to-executable-plan completion. The method unifies natural language understanding, formal modeling, and robot simulation, and is validated in dynamic environments. Experiments demonstrate significant improvements over baseline planners in both valid plan generation rate and task success rate, alongside enhanced environmental adaptability and robustness. The core contribution lies in establishing a verifiable and interpretable synergy between LLM-based commonsense reasoning and symbolic planning—bridging neural and symbolic AI in a principled, transparent manner.
This work addresses the performance bottleneck posed by full grounding in classical planning, which suffers from exponential growth in the number of actions and atoms in large-scale tasks. To overcome this limitation, the paper introduces large language models (LLMs) into the partial grounding process for the first time. By leveraging the semantic structure of PDDL domain and problem files, the approach heuristically identifies and prunes irrelevant objects, actions, and predicates, enabling efficient partial grounding. This method transcends the constraints of traditional techniques based on dependency graphs or embeddings, achieving speedups of multiple orders of magnitude on seven challenging grounding benchmarks while maintaining comparable or even superior plan quality across several domains.
This work addresses the critical limitation of existing vision-language models in task planning, which often neglect spatial executability and thus fail to guide real-world robotic manipulation. To bridge this gap, we introduce a novel task termed "spatially grounded long-horizon task planning," establish a benchmark dataset named GroundedPlanBench, and propose the V2GP framework. V2GP leverages real robot demonstration videos to automatically generate hierarchical planning data annotated with spatial grounding, enabling joint optimization of high-level action sequences and low-level spatial interaction points. Experimental results demonstrate that V2GP significantly enhances the spatial executability of generated plans on both GroundedPlanBench and physical robot platforms, advancing task planning toward practical deployment in real-world environments.
Long-horizon hierarchical planning in text-based environments faces challenges including open-ended action spaces, ambiguous observations, and sparse rewards; existing LLM-dependent approaches suffer from high inference overhead, non-differentiable parameters, and inefficient deployment. Method: We propose the “One-Shot Teacher” paradigm: an LLM is invoked only once at planning initialization to generate a subgoal sequence, followed by LLM-guided trajectory distillation to pretrain a lightweight student planner (e.g., a Transformer-based planner) for subgoal-conditioned modeling. Contribution/Results: This eliminates repeated LLM calls during both training and inference. On TextCraft, our method achieves 56% success rate—surpassing ADaPT’s 52%—while reducing average inference time from 164.4 seconds to 3.0 seconds (54× speedup), significantly improving both task performance and deployment efficiency.