Embodied Chain of Action Reasoning with Multi-Modal Foundation Model for Humanoid Loco-manipulation

πŸ“… 2025-04-13
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Humanoid robots struggle with long-horizon, multimodal coordination of locomotion-manipulation tasks in complex, unstructured environments. Method: This paper proposes the Embodied Chain-of-Action (ECOA) frameworkβ€”a novel, humanoid-specific chain-of-thought paradigm grounded in multimodal foundation models. ECOA jointly models object affordances, whole-body kinematics, and spatial reasoning to enable cross-modal action generation under occlusion or for unseen objects; it further employs upper-lower limb decoupled control and joint vision-language-action representation learning to achieve end-to-end mapping from natural language instructions to coordinated whole-body actions. Contribution/Results: Evaluated on real-world object rearrangement and loco-manipulation tasks, ECOA significantly improves instruction understanding accuracy and task success rate, effectively bridging the semantic gap between high-level planning and low-level execution.

Technology Category

Intelligent Robots: Embodied AIHumans and AI: Human-Aware Planning and Behavior PredictionMachine Learning: Large Multimodal Models (LMMs)

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAG
πŸ“ Abstract
Enabling humanoid robots to autonomously perform loco-manipulation tasks in complex, unstructured environments poses significant challenges. This entails equipping robots with the capability to plan actions over extended horizons while leveraging multi-modality to bridge gaps between high-level planning and actual task execution. Recent advancements in multi-modal foundation models have showcased substantial potential in enhancing planning and reasoning abilities, particularly in the comprehension and processing of semantic information for robotic control tasks. In this paper, we introduce a novel framework based on foundation models that applies the embodied chain of action reasoning methodology to autonomously plan actions from textual instructions for humanoid loco-manipulation. Our method integrates humanoid-specific chain of thought methodology, including detailed affordance and body movement analysis, which provides a breakdown of the task into a sequence of locomotion and manipulation actions. Moreover, we incorporate spatial reasoning based on the observation and target object properties to effectively navigate where target position may be unseen or occluded. Through rigorous experimental setups on object rearrangement, manipulations and loco-manipulation tasks on a real-world environment, we evaluate our method's efficacy on the decoupled upper and lower body control and demonstrate the effectiveness of the chain of robotic action reasoning strategies in comprehending human instructions.
Problem

Research questions and friction points this paper is trying to address.

Enabling humanoid robots to autonomously perform loco-manipulation in complex environments
Bridging gaps between high-level planning and task execution using multi-modality
Planning actions from textual instructions with embodied chain of reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-modal foundation model for action planning
Embodied chain of action reasoning methodology
Spatial reasoning for unseen target navigation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Y
Yu Hao
Embodied AI and Robotics Lab, New York University Abu Dhabi, UAE
G
Geeta Chandra Raju Bethala
Embodied AI and Robotics Lab, New York University Abu Dhabi, UAE
N
Niraj Pudasaini
Embodied AI and Robotics Lab, New York University Abu Dhabi, UAE
H
Hao Huang
Embodied AI and Robotics Lab, New York University Abu Dhabi, UAE
S
Shuaihang Yuan
Embodied AI and Robotics Lab, New York University Abu Dhabi, UAE
C
Congcong Wen
Embodied AI and Robotics Lab, New York University Abu Dhabi, UAE
Baoru Huang
Baoru Huang
University of Liverpool; Imperial College London
RoboticsComputer visionSurgical visionImage-Guided Intervention
A
Anh Nguyen
Department of Computer Science, University of Liverpool, UK
Y
Yi Fang
Embodied AI and Robotics Lab, New York University Abu Dhabi, UAE