🤖 AI Summary
This study addresses the lack of physical interactivity and task repeatability in robotic training data by proposing a graph-native coordinated agent framework. The framework integrates ISArt component generation with dual pathways—Unravel for image reconstruction and SimForge for text generation—to enable automatic, task-bound synthesis of editable, physically valid simulation environments from images or text, while supporting module-level fault repair and state preservation. A pipeline built upon MuJoCo and Isaac Sim significantly enhances policy learning performance. Evaluated on the LIBERO-Plus benchmark, the proposed approach increases the success rate of world action models from 77.17% to 89.43%, thereby validating the effectiveness of cross-simulator policy transfer.
📝 Abstract
Generating useful robot-training data requires more than visually plausiblescenes: objects must support interaction, placements must remain physicallyvalid, and tasks must admit repeatable execution. We present\textbf{Awomo-SimDataEngine}, an agentic system that connects asset and scenegeneration to robot demonstration synthesis. Shared asset services providerigid and articulated objects, including structure-grounded part and jointgeneration with ISArt. Scene generation supports two complementary routes:Unravel reconstructs editable scenes from images, while SimForge buildssingle-room and multi-room environments from text. A graph-native harnesscoordinates construction, validation, andbounded repair, routing failures to the responsible module while retainingunaffected scene state. PolicyForge binds validated worlds to tasks and robotembodiments to produce replayable demonstrations. Evaluations cover assetgeometry, scene quality, and downstream policy learning. On MuJoCo-basedLIBERO-Plus, co-training with Isaac Sim demonstrations improves the overallsuccess rate of a World-Action Model (WAM) from $77.17\%$ to $89.43\%$. Goal and spatialsuccess improve by $31.66$ and $6.25$ percentage points, respectively.These results support the utility of the generated data for cross-simulatorpolicy training, with more limited gains on long-horizon tasks.