Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation

📅 2025-01-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Achieving generalizable 3D manipulation in dynamic environments remains challenging for robotic systems. Method: This paper introduces LMM-3DP, the first framework to tightly integrate high-level semantic planning from Large Multimodal Models (LMMs) with low-level, semantics-aware 3D feature field control. Contribution/Results: Its core innovations include (1) a language-3D joint attention mechanism enabling cross-modal feature alignment within a 3D Transformer; (2) a closed-loop collaborative architecture incorporating self-feedback critique, hierarchical policy memory, and failure-driven retry; and (3) robust long-horizon kitchen manipulation under environmental perturbations. Experiments in real kitchen settings demonstrate a 1.45× improvement in low-level control success rate and ~1.5× gain in high-level planning accuracy over pure LLM-based baselines, significantly advancing embodied 3D reasoning and execution.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Intelligent Robots: ManipulationNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
The recent advancements in visual reasoning capabilities of large multimodal models (LMMs) and the semantic enrichment of 3D feature fields have expanded the horizons of robotic capabilities. These developments hold significant potential for bridging the gap between high-level reasoning from LMMs and low-level control policies utilizing 3D feature fields. In this work, we introduce LMM-3DP, a framework that can integrate LMM planners and 3D skill Policies. Our approach consists of three key perspectives: high-level planning, low-level control, and effective integration. For high-level planning, LMM-3DP supports dynamic scene understanding for environment disturbances, a critic agent with self-feedback, history policy memorization, and reattempts after failures. For low-level control, LMM-3DP utilizes a semantic-aware 3D feature field for accurate manipulation. In aligning high-level and low-level control for robot actions, language embeddings representing the high-level policy are jointly attended with the 3D feature field in the 3D transformer for seamless integration. We extensively evaluate our approach across multiple skills and long-horizon tasks in a real-world kitchen environment. Our results show a significant 1.45x success rate increase in low-level control and an approximate 1.5x improvement in high-level planning accuracy compared to LLM-based baselines. Demo videos and an overview of LMM-3DP are available at https://lmm-3dp-release.github.io.
Problem

Research questions and friction points this paper is trying to address.

Robotics
Adaptive Control
3D Spatial Manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrated LMM Planner
3D Skill Policies
Enhanced Cognitive Capabilities
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuelei Li
UC San Diego
G
Ge Yan
UC San Diego
A
Annabella Macaluso
UC San Diego
M
Mazeyu Ji
UC San Diego
Xueyan Zou
Xueyan Zou
PostDoc at UC San Diego
Foundation ModelRobotic
X
Xiaolong Wang
UC San Diego