Data Pyramid for Embodied Manipulation

šŸ“… 2026-07-27
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
This work addresses the lack of a systematic organizational framework for embodied intelligence data, which hinders scalability and robot alignment. The authors propose a ā€œData Pyramidā€ framework that introduces a hierarchical structure to integrate five heterogeneous data sources—real robot data, UMI-style datasets, first- and third-person videos, simulation data, and general vision-language corpora—organized according to quality, diversity, reusability, and physical fidelity. Through multimodal data evaluation, alignment strategies, and hybrid pretraining, the study systematically analyzes how data composition influences model perception, reasoning, and planning capabilities. The work establishes design principles for data selection and combination in embodied foundation models and identifies six key open challenges, thereby advancing the development of data infrastructure for embodied learning.
šŸ“ Abstract
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.
Problem

Research questions and friction points this paper is trying to address.

embodied manipulation
multimodal data
data pyramid
robot learning
foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Pyramid
Embodied Manipulation
Multimodal Foundation Models
Robot Learning
Data Composition
šŸ”Ž Similar Papers
Y
Yifan Ye
PKU
Y
Yankai Fu
PKU
Y
Yaoxu Lv
PKU
Bohan Hou
Bohan Hou
PhD of Computer Science, Carnegie Mellon University
Machine LearningSystems
J
Jun Cen
HKUST
Lingdong Kong
Lingdong Kong
National University of Singapore
Computer VisionDeep Learning
Duo Zheng
Duo Zheng
The Chinese University of Hong Kong
Computer Vision
T
Tianxing Chen
HKU
J
Jiaming Liu
PKU
Ziang Cao
Ziang Cao
Nanyang Technological University
Deep learningRobotics
Y
Yunfan Lou
PKU
Wei Chow
Wei Chow
Zhejiang University
vision-languagegenerative AI
Xian Sun
Xian Sun
AerospaceĀ InformationĀ ResearchĀ Institute,Ā ChineseĀ AcademyĀ ofĀ Sciences
Remote SensingComputer Vision and Pattern RecognitionArtificial Intelligence
Y
Yingshuo Wang
UCB
Kuangzhi Ge
Kuangzhi Ge
Peking University
Multimodal LearningEmbodied AI
Xiaowei Chi
Xiaowei Chi
The Hong Kong University of Science and Technology
Multimodal GenerationRoboticsComputer Vision
X
Xidong Zhang
GBU
Zhibo Pang
Zhibo Pang
ABB Corporate Research, and KTH Royal Institute of Technology, Sweden
RoboticsAICloudWirelessIndustrial Automation
Yiwu Zhong
Yiwu Zhong
CUHK / University of Wisconsin-Madison
Vision-Language LearningMulti-Modal ModelsEmbodied AI
Sirui Han
Sirui Han
The Hong Kong University of Science and Technology
Large Language ModelInterdisciplinary Artificial Intelligence
Zhihe Lu
Zhihe Lu
HBKU<--NUS<--University of Surrey<--CASIA
Computer VisionTransfer LearningFew-shot LearningMultimodel LearningContinual Learning
Weihao Yuan
Weihao Yuan
Hong Kong University of Science and Technology
3D VisionEmbodied AIRobot Reinforcement Learning
Qifeng Chen
Qifeng Chen
HKUST
Computational PhotographyImage SynthesisGenerative AIAutonomous DrivingEmbodied AI
Michael Yu Wang
Michael Yu Wang
Chair Professor & Dean, Great Bay University, China
RoboticsTopology OptimizationAdditive Manufacturing
Y
Yao Mu
SJTU