Interactive Task Planning with Language Models

๐Ÿ“… 2023-10-16
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 24
โœจ Influential: 2
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing interactive robotic task planners suffer from poor generalization, reliance on predefined modules, and labor-intensive prompt engineering. Method: This paper proposes a long-horizon task planning framework for real-world scenarios, introducing a novel functional architecture that jointly integrates high-level language-based planning with low-level skill invocation, grounded in vision-language multimodal scene understanding. Contribution/Results: The framework enables dynamic goal interpretation, online replanning, and zero-shot cross-task transferโ€”requiring only lightweight task instructions without domain-specific fine-tuning or manual prompt design. Evaluated on a real-world bubble tea preparation task, it successfully generates and executes unseen goals, accommodates mid-task user requests for modification, performs precise online replanning, and generalizes to diverse household service scenarios.
๐Ÿ“ Abstract
An interactive robot framework accomplishes long-horizon task planning and can easily generalize to new goals and distinct tasks, even during execution. However, most traditional methods require predefined module design, making it hard to generalize to different goals. Recent large language model based approaches can allow for more open-ended planning but often require heavy prompt engineering or domain specific pretrained models. To tackle this, we propose a simple framework that achieves interactive task planning with language models by incorporating both high-level planning and low-level skill execution through function calling, leveraging pretrained vision models to ground the scene in language. We verify the robustness of our system on the real world task of making milk tea drinks. Our system is able to generate novel high-level instructions for unseen objectives and successfully accomplishes user tasks. Furthermore, when the user sends a new request, our system is able to replan accordingly with precision based on the new request, task guidelines and previously executed steps. Our approach is easy to adapt to different tasks by simply substituting the task guidelines, without the need for additional complex prompt engineering. Please check more details on our https://wuphilipp.github.io/itp_site and https://youtu.be/TrKLuyv26_g.
Problem

Research questions and friction points this paper is trying to address.

Interactive robot task planning generalization
Language models for open-ended planning
Adapting to new tasks without complex engineering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interactive task planning framework
Function calling for skill execution
Pretrained vision models integration
๐Ÿ”Ž Similar Papers
No similar papers found.
University of California, Berkeley
B
Boyi Li
University of California, Berkeley
P
Philipp Wu
University of California, Berkeley
Pieter Abbeel
Pieter Abbeel
UC Berkeley | Covariant
RoboticsMachine LearningAI
J
Jitendra Malik
University of California, Berkeley