๐ค AI Summary
This work addresses the limited capability of service and assistant robots in task and motion planning during natural language interaction by proposing a hierarchical language-driven framework that decouples high-level task planning from low-level spatial reasoning through the collaboration of two large language models (LLMs). The high-level agent interprets natural language instructions to generate action sequences, while the low-level module integrates YOLOX-GDRNet for object detection and pose estimation, employing ReAct-style prompting and tool-calling mechanisms to handle 3D spatial placement tasks and identify infeasible requests. Evaluated across 24 test scenarios ranging from simple to complex instructions, the system achieves an end-to-end task success rate of 86%, significantly enhancing the intuitiveness and robustness of human-robot collaboration.
๐ Abstract
We present a hierarchical language-driven framework for robotic task and motion planning to improve natural, intuitive human-robot interaction in service and assistance scenarios. The proposed system employs two large language model (LLM) modules: a high-level planning agent and a low-level spatial reasoning sub-module. The primary agent processes natural language commands and generates action sequences using a ReAct-style prompt, interacting with tools for object perception and manipulation (e.g., pick, place, release). For precise spatial placement, such as interpreting "place the mug next to the plate", a separate sub-prompting module handles 3D reasoning based on object geometry and scene layout. The system integrates YOLOX-GDRNet for object detection and pose estimation, along with a motion execution stub. We evaluated the system in 24 test scenarios, ranging from simple spatial commands to high-level instructions and infeasible requests. The system achieved an overall task success rate of 86%.