language-grounded planning

Designs and builds planners that translate natural-language or multimodal instructions into executable action sequences or compact latent plan representations encoding semantic intent and goal conditions. These systems ground language to perceived objects, spatial relations, and targets, produce instruction rewrites or grounding cues, and incorporate constraints (e.g., accessibility, social norms) using techniques such as latent reasoning tokens and mutual-information objectives for goal-conditioned semantic planning.

language-groundedplanning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Grounding Language Models with Semantic Digital Twins for Robotic Planning

Jun 19, 2025
MN
Mehreen Naeem
🏛️ University of Bremen

This work addresses the challenges of adaptability and robustness in robotic execution of natural language instructions within dynamic environments. We propose a deep synergy framework integrating large language models (LLMs) with semantic digital twins (SDTs). Methodologically, we introduce the first bidirectional coupling: the SDT provides real-time, semantically grounded environmental representations and affordance-aware perception, while the LLM performs instruction parsing, reflective reasoning, and generation of structured action triplets—augmented with failure-driven iterative replanning. Our technical contributions include SDT modeling, adaptation to the ALFRED benchmark, and closed-loop simulation evaluation. Experiments demonstrate significant improvements in task success rate and fault tolerance across diverse household scenarios, effectively handling uncertainties such as object pose variations, occlusions, and execution failures. The framework establishes a novel paradigm for semantic-level task planning in embodied intelligence.

Decompose language instructions into structured action tripletsGenerate recovery strategies for execution failuresIntegrate SDTs and LLMs for robotic task execution

Plan-X: Instruct Video Generation via Semantic Planning

Nov 22, 2025
LH
Lun Huang
🏛️ Duke University | ByteDance Intelligent Creation | Princeton University

Diffusion Transformers often suffer from visual hallucinations and poor instruction alignment in complex video generation—particularly for high-level semantic tasks involving human-object interaction, multi-stage actions, and contextual motion reasoning. To address these limitations, we propose Plan-X, the first framework to introduce a learnable multimodal semantic planning mechanism. Plan-X employs a multimodal large language model as a semantic planner that jointly reasons over textual and visual context to infer user intent and autoregressively generate spatiotemporal semantic tokens. These tokens serve as structured priors that explicitly guide the diffusion Transformer toward high-fidelity video synthesis. Experiments demonstrate that Plan-X significantly suppresses visual hallucinations, improves instruction adherence and spatiotemporal consistency, and achieves state-of-the-art performance on complex scene understanding and long-horizon action modeling tasks.

Address visual hallucinations in video generation modelsEnhance semantic reasoning for multi-stage actionsImprove alignment with complex user instructions

Current large model–based approaches to autonomous driving scene understanding and planning lack effective temporal modeling, leading to inconsistent reasoning over sequential actions and compromising both safety and interpretability. To address this, this work proposes three multi-agent planner architectures incorporating varying degrees of temporal conditioning constraints. The authors establish the first empirical benchmark for temporally aware scene-to-planning reasoning on a subset of BDD-X and introduce evaluation metrics assessing semantic, syntactic, and logical consistency. Experimental results show that while explicit temporal constraints do not significantly improve standard NLP metrics, qualitative analysis reveals their capacity to elicit forward-looking risk assessment, stabilize corrective behaviors, and enhance strategic diversity. The study also highlights limitations in current prompt engineering practices regarding temporal grounding.

agent communicationautonomous vehiclesscene-to-plan reasoning

Socratic Planner: Inquiry-Based Zero-Shot Planning for Embodied Instruction Following

Apr 21, 2024
SS
Suyeon Shin
🏛️ Seoul National University

To address the challenges of zero-shot planning difficulty and weak dynamic adaptability in Embodied Instruction Following (EIF), this paper proposes the first training-free, vision-driven zero-shot embodied planning framework. Methodologically, it decomposes natural language instructions into executable high-level sub-goal sequences via a self-questioning-and-answering mechanism; introduces a visually grounded real-time re-planning mechanism that dynamically refines plans based on visual feedback during interaction; and designs RelaxedHLP—a novel evaluation metric that quantifies high-level planning quality for the first time. Experiments on the ALFRED benchmark demonstrate state-of-the-art zero-shot and few-shot performance, particularly on complex tasks requiring multi-step reasoning and environment responsiveness. Visual feedback significantly improves re-planning accuracy, with consistent gains over existing methods.

Handling unexpected obstacles during task executionImproving long-horizon task performance with complex inferenceZero-shot planning for embodied instruction following tasks

Existing robotic navigation systems struggle with unstructured, map-free environments when receiving incomplete natural-language task instructions. Method: This work proposes an online semantic planning framework leveraging large language models (LLMs), integrated into a closed-loop system that synergistically combines real-time semantic SLAM, receding-horizon planning (RHP), and online safety verification. The framework enables concurrent semantic mapping, automatic task completion, dynamic subtask re-planning, and runtime safety constraint enforcement—without requiring prior maps or manually refined commands. Contribution/Results: It establishes the first end-to-end online pipeline for semantic understanding, planning, and execution. Evaluated in cluttered outdoor environments exceeding 20,000 m², the approach reduces task completion time and path length by over 50% compared to baselines, while significantly decreasing user interaction frequency.

Enhancing efficiency and reducing user interaction in robotic missionsMapping and planning in unstructured environments without pre-built mapsOnline semantic planning for incomplete natural language missions

Latest Papers

What's happening recently
View more

Robots frequently fail to execute everyday tasks due to natural language instructions lacking commonsense preconditions and decomposed subgoals. This paper proposes an LLM-augmented symbolic planning framework that, for the first time, automatically formalizes implicit preconditions and refined subgoals generated by large language models (LLMs) and integrates them into classical planners (e.g., PDDL), enabling end-to-end instruction-to-executable-plan completion. The method unifies natural language understanding, formal modeling, and robot simulation, and is validated in dynamic environments. Experiments demonstrate significant improvements over baseline planners in both valid plan generation rate and task success rate, alongside enhanced environmental adaptability and robustness. The core contribution lies in establishing a verifiable and interpretable synergy between LLM-based commonsense reasoning and symbolic planning—bridging neural and symbolic AI in a principled, transparent manner.

LLMs generate preconditions and subgoals to improve planning reliabilityRobots fail tasks due to missing commonsense detailsTraditional planners require explicit, time-consuming manual specification

This work addresses the performance bottleneck posed by full grounding in classical planning, which suffers from exponential growth in the number of actions and atoms in large-scale tasks. To overcome this limitation, the paper introduces large language models (LLMs) into the partial grounding process for the first time. By leveraging the semantic structure of PDDL domain and problem files, the approach heuristically identifies and prunes irrelevant objects, actions, and predicates, enabling efficient partial grounding. This method transcends the constraints of traditional techniques based on dependency graphs or embeddings, achieving speedups of multiple orders of magnitude on seven challenging grounding benchmarks while maintaining comparable or even superior plan quality across several domains.

classical planningcomputational bottleneckgrounding

This work addresses the critical limitation of existing vision-language models in task planning, which often neglect spatial executability and thus fail to guide real-world robotic manipulation. To bridge this gap, we introduce a novel task termed "spatially grounded long-horizon task planning," establish a benchmark dataset named GroundedPlanBench, and propose the V2GP framework. V2GP leverages real robot demonstration videos to automatically generate hierarchical planning data annotated with spatial grounding, enabling joint optimization of high-level action sequences and low-level spatial interaction points. Experimental results demonstrate that V2GP significantly enhances the spatial executability of generated plans on both GroundedPlanBench and physical robot platforms, advancing task planning toward practical deployment in real-world environments.

action planninglong-horizon planningrobot manipulation

Long-horizon hierarchical planning in text-based environments faces challenges including open-ended action spaces, ambiguous observations, and sparse rewards; existing LLM-dependent approaches suffer from high inference overhead, non-differentiable parameters, and inefficient deployment. Method: We propose the “One-Shot Teacher” paradigm: an LLM is invoked only once at planning initialization to generate a subgoal sequence, followed by LLM-guided trajectory distillation to pretrain a lightweight student planner (e.g., a Transformer-based planner) for subgoal-conditioned modeling. Contribution/Results: This eliminates repeated LLM calls during both training and inference. On TextCraft, our method achieves 56% success rate—surpassing ADaPT’s 52%—while reducing average inference time from 164.4 seconds to 3.0 seconds (54× speedup), significantly improving both task performance and deployment efficiency.

Addresses high computational cost of LLM-based hierarchical planningEliminates need for repeated LLM queries during training and inferenceUses one-shot LLM guidance to pretrain a lightweight student model

Hot Scholars

CA

Christopher Agia

PhD Student in Computer Science, Stanford University
RoboticsMachine LearningComputer VisionArtificial Intelligence
TH

Tucker Hermans

Associate Professor, School of Computing, University of Utah; Senior Research Scientist, NVIDIA
RoboticsManipulationRobot LearningMotion Planning
WL

Weiwen Liu

Associate Professor, Shanghai Jiao Tong University
large language modelsAI agentsrecommender systems
WH

William Hunt

Postgraduate Researcher, University of Southampton