š¤ AI Summary
This work addresses the lack of a systematic organizational framework for embodied intelligence data, which hinders scalability and robot alignment. The authors propose a āData Pyramidā framework that introduces a hierarchical structure to integrate five heterogeneous data sourcesāreal robot data, UMI-style datasets, first- and third-person videos, simulation data, and general vision-language corporaāorganized according to quality, diversity, reusability, and physical fidelity. Through multimodal data evaluation, alignment strategies, and hybrid pretraining, the study systematically analyzes how data composition influences model perception, reasoning, and planning capabilities. The work establishes design principles for data selection and combination in embodied foundation models and identifies six key open challenges, thereby advancing the development of data infrastructure for embodied learning.
š Abstract
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.