Score
Designs and implements systems that use large language models to translate textual or semantic instructions into time-indexed motion trajectories, waypoints, or control commands for agents, cameras, or objects. Builds and evaluates pipelines that ensure feasibility, smoothness, collision- and constraint-aware execution, and temporal coordination or synchronization across multiple actors or sensors.
Real-time robotic trajectory adaptation to dynamic human instructions remains challenging due to poor generalization and limited interpretability. Method: We propose a language-grounded trajectory adjustment framework that directly leverages pre-trained large language models (LLMs) to generate executable code-based policies—enabling semantic understanding and numerically grounded refinement of planner outputs or demonstration trajectories, without task-specific fine-tuning. Our approach integrates code-generation–driven policy modeling, multi-platform simulation (PyBullet/Gazebo), generic motion planners (RRT/A*), and human demonstration transfer. Results: Evaluated on simulated robotic arms, UAVs, and ground robots, the framework achieves high-precision trajectory adaptation to complex, numerically parameterized multi-step instructions (e.g., “detour 0.3 m around the right-side obstacle and decelerate to stop”). It significantly outperforms existing feature-driven sequential models, demonstrating strong generalization, inherent interpretability, and natural interactive feedback capability.
This work addresses the challenge of precisely controlling object dynamics and camera trajectories in video generation via natural language—a longstanding limitation in semantic video synthesis. To this end, we propose the first end-to-end language-driven framework for joint 3D object–camera motion generation. Methodologically, we design a domain-specific language (DSL) tailored to cinematic motion, leveraging large language models and program synthesis to automatically parse natural language descriptions into structured, executable 3D trajectory programs. We further introduce the first large-scale text–program–trajectory triplet dataset to support training and evaluation. Compared to prior approaches, our method significantly improves motion controllability and alignment with user intent, while preserving high-fidelity 3D motion planning. Crucially, it offers strong interpretability and post-hoc editability of generated motions. This establishes a novel paradigm for semantic, film-grade video generation.
This work addresses the challenge of executing complex semantic tasks in large-scale outdoor multi-robot systems via natural language instructions. We propose an integrated 3D scene graph framework that unifies multi-robot SLAM, open-vocabulary object detection and mapping, LLM-driven semantic parsing, and hierarchical Task-and-Motion Planning (TAMP). We introduce a novel language-guided PDDL goal generation mechanism to close the loop from intent to executable actions, and design a view-invariant relocalization method with shared scene graph fusion for collaborative multi-robot operation. Evaluated in real-world large-scale outdoor environments, our system achieves significant improvements in natural language understanding accuracy and task success rate, reduces relocalization error by 42%, and maintains planning response latency under 800 ms. The core contribution is the first end-to-end 3D semantic planning system supporting open-set object recognition, language-guided goal generation, and multi-robot collaborative relocalization.
研究通过约束大语言模型和物理验证,解决机器人在复杂环境中安全执行任务的问题,提高操作成功率。
This work addresses the critical challenge of generating safe, feasible, and context-aware interactive motion trajectories for autonomous driving conditioned on natural language instructions. Methodologically, we propose a semantic-guided multimodal motion prediction framework centered on the first text-instruction-driven multimodal large language model (MLLM), integrating a pretrained LLM, LoRA-based efficient fine-tuning, and multimodal scene encoding—trained on our newly curated InstructWaymo dataset. Crucially, we introduce the first instruction feasibility identification module with an active rejection mechanism to handle infeasible commands. Experiments on the Waymo Open Motion Dataset demonstrate that our model achieves high trajectory generation accuracy for feasible instructions and significantly outperforms baselines in rejecting infeasible ones. These results validate the effectiveness of semantic guidance in enhancing dynamic scene understanding and safety-critical response capabilities.
This work addresses the challenge that existing vision-language-action models struggle to accurately execute natural language instructions involving spatiotemporal and logical constraints, while also lacking interpretability. The authors propose a hierarchical framework that, for the first time, deeply integrates Signal Temporal Logic (STL) between language understanding and robotic execution. The approach decomposes high-level instructions into subtasks and generates verifiable, optimizable, and correctable STL specifications, which dynamically schedule low-level policies. By combining vision-language models, STL, model predictive control, and learned policies, the method enables an end-to-end mapping from natural language instructions to formal specifications, supporting online monitoring and replanning. Experiments in real-world tabletop environments demonstrate significant improvements in accuracy, reliability, and interpretability of language-guided robotic tasks.
This work proposes an end-to-end approach leveraging large language models to automatically translate natural language descriptions of space mission intent into solvable trajectory optimization formulations, substantially reducing reliance on domain experts. It represents the first application of large language models to the semantic-to-formal modeling pipeline in spacecraft trajectory optimization, integrating natural language understanding, convex optimization, and spacecraft rendezvous dynamics to generate executable optimization code directly from high-level mission specifications. Evaluated in spacecraft rendezvous scenarios, the system demonstrates a high success rate in reconstructing feasible convex optimization problems, thereby validating its effectiveness, practical utility, and significant enhancement of both mission design flexibility and development efficiency.
Existing vision-language-action (VLA) models struggle to simultaneously achieve effective global planning and diverse physical manipulation in long-horizon tasks. This work proposes the VLAs-as-Tools framework, which employs a high-level vision-language model for global planning and recovery, while delegating local subtasks to multiple specialized VLA tools. A unified tool interface enables event-triggered replanning, tightly coupling planning and execution. The approach introduces a novel VLA tool-family interface and a Tool-Aligned Post-Training (TAPT) method to enhance tool-call fidelity and coordination. Evaluated on LIBERO-Long and RoboTwin, the framework improves task success rates by 4.8% and 23.1%, respectively, and increases tool-call fidelity by 15.0%.
This work addresses the challenge of reliably and interpretably translating open-ended natural language instructions from passengers into low-level control signals for autonomous vehicles, while ensuring real-time performance and safety. The authors propose a scheduler-centric execution framework that leverages a large language model (LLM) to parse semantic instructions and generate executable scripts, which dynamically orchestrate multiple model predictive control (MPC)-based motion planners. A closed-loop feedback mechanism enables real-time generation of control signals, establishing a transparent decision chain from semantics to actions. By integrating LLMs with a multi-planner scheduling architecture, this approach decouples the timescales of semantic reasoning and vehicle control, yielding a traceable and interpretable system. The study also introduces the first high-fidelity benchmark for evaluating open-ended instruction following. Experiments demonstrate significant improvements in task completion rates, reduced LLM query costs, safety and compliance on par with specialized methods, and strong robustness to LLM inference latency.
This study addresses the challenge of translating natural language intents into robot-executable actions within dynamic, unknown environments by proposing an intent-driven dual-AI collaborative framework. The framework leverages large language models to generate constrained executable code and integrates vision-language models for semantic grounding. Its core innovation lies in an adaptive replanning mechanism triggered by geometric and semantic thresholds, which achieves closed-loop control through runtime monitoring. Experimental results demonstrate that the proposed approach robustly executes complex instructions under bounded reaction cycles, enabling effective robot control in dynamic settings. These findings validate the reliability and generalization capability of generative AI for real-time embodied intelligence tasks.