🤖 AI Summary
This study addresses the scarcity of training data for robot foundation models and the limitations of existing simulation pipelines in coordinating assets, scenes, and tasks. We propose an end-to-end automated data generation framework driven by a recursive self-improving flywheel. This approach introduces a novel joint iterative optimization mechanism for scenes and tasks, transcending predefined template constraints. By integrating agent-in-the-loop refinement cycles, language-driven customization, and high-fidelity physics simulation, the framework supports complex embodied interactions involving mobile manipulation, humanoid robots, and fluid dynamics. Experimental results demonstrate that our method significantly improves generation success rates for long-horizon tasks. Furthermore, we validate that the diversity of the generated data substantially enhances the generalization capabilities of downstream policies.
📝 Abstract
Scaling robotic foundation models requires diverse training data and reliable evaluation environments. Simulation offers a scalable solution, yet existing generation pipelines remain constrained by predefined assets and skills, a disconnect between scene generation and task generation, and limited support for complex embodiments and physics. We introduce EmbodiedSmith, a framework for scalable embodied data generation through recursive self-improvement (RSI). EmbodiedSmith unifies asset, scene, and task generation in a pipeline that supports autonomous creation and language-driven customization. Its core is an agentic refinement loop: scene generation anticipates downstream task requirements, while task generation guides targeted scene edits, allowing scenes and tasks to iteratively improve one another. This joint refinement improves task generation success, including for long-horizon tasks. The framework further supports mobile manipulators, humanoids, and dexterous hands, as well as interactions involving deformable objects and fluids, broadening the range of behaviors and physical phenomena represented in generated data. Together, these capabilities provide a flexible simulation engine for both robot pretraining and evaluation. Extensive experiments validate the quality, diversity, and generation efficiency of the resulting data, while downstream policy experiments demonstrate that increased data diversity improves generalization.