MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究通过Mixture-of-Transformers架构的MachEmbodied-U0模型,解决了通用机器人控制中理解任务意图与生成精确动作的问题,结合视觉动态和动作生成以提高操作精度。
📝 Abstract
General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual prediction with control without necessarily exposing the task-relevant semantic and spatial structure needed for fine-grained manipulation. We present MachEmbodied-U0 (ME-U0), a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture. Subtask prediction and affordance grounding guide joint visual-dynamics and action generation via flow matching. Visual dynamics encompass future RGB, depth, surface normals, and optical flow, providing complementary supervision for appearance, geometry, and motion. Multi-rate Rotary Position Encoding (MRPE) aligns visual dynamics with fine-grained control. We pretrain ME-U0 on approximately 4,200 hours of curated demonstrations from robotic datasets and egocentric datasets. Using only the supervision natively available in each downstream benchmark, ME-U0 achieves an average score of 17.66 on the RoboDojo simulation benchmark and average success rates of 99.0\% and 82.5\% on LIBERO and LIBERO-Plus, respectively. We additionally validate ME-U0 on real-world robotic manipulation tasks, demonstrating its effectiveness beyond simulation. Without corresponding downstream supervision, ME-U0 further demonstrates zero-shot subtask prediction, affordance grounding, and visual dynamics on simulated and real-world observations. Overall, ME-U0 combines competitive downstream control performance with transferable task-grounding and visual-dynamics capabilities across simulation and the real world.
Problem

Research questions and friction points this paper is trying to address.

robot control
task intent
scene dynamics
semantic priors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Transformers
flow matching
Multi-rate Rotary Position Encoding (MRPE)
visual dynamics
zero-shot
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haoran Wen
Li Auto Inc.
W
Wenfu Wang
Li Auto Inc.
K
Kunsong Shi
Li Auto Inc.
J
Jingke Wang
Li Auto Inc.
Wancheng Feng
Wancheng Feng
Student of Shandong University of Science and Technology
Computer VisionAIGCGenerative AI3D
Y
Yiren Zhang
Li Auto Inc.
Y
Yueran Zhao
Li Auto Inc.
Xuancheng Zhang
Xuancheng Zhang
Tsinghua University
3D VisionDeep Learning
N
Nanfei Ye
Li Auto Inc.
X
Xingru Chen
Li Auto Inc.
Zhaohong Sun
Zhaohong Sun
Associate Professor@Kyushu University, Research Scientist@CyberAgent
Artificial IntelligenceAlgorithmic Mechanism DesignMatching Theory
C
Chengmin Yang
Li Auto Inc.
Z
Zikang Yu
Li Auto Inc.
P
Penghao Bi
Li Auto Inc.
J
Jia Shi
Li Auto Inc.
Y
Yu Liu
Li Auto Inc.
K
Kun Zhan
Li Auto Inc.
Y
Yan Xie
Li Auto Inc.