🤖 AI Summary
This study addresses the lack of large-scale, verifiable training environments for embodied intelligence and the difficulty in reproducing the data scaling effects observed in vision-language models. To this end, it proposes EnvDreamer, a framework that pioneers an automated paradigm for generating high-quality interactive environments from multimodal instructions. By integrating large language and multimodal models for semantic parsing with Unreal Engine 5 rendering, and employing automatic verifiers to ensure environmental compliance, the approach enables large-scale policy training without explicit maps or human supervision. Experiments demonstrate that agents trained in these generated environments achieve competitive results across multiple benchmarks. Furthermore, the open-sourced EnvDreamer-20k dataset bridges a critical gap in world model data, providing essential infrastructure for reproducible evaluation in embodied AI research.
📝 Abstract
Large datasets and high capacity models have accelerated progress in vision and language. This work introduces a platform aimed at bringing comparable gains to embodied learning, world models, and robotics. We present EnvDreamer, a framework that uses large language and vision language models to generate Unreal Engine 5 environments for embodied AI and robot training. EnvDreamer enables sampling of large, diverse, interactive, customizable, and validator passed virtual environments for training and evaluation across navigation, interaction, and manipulation. We illustrate the platform with a large set of generated scenes and simple baselines. Policies trained on EnvDreamer generated environments, without explicit mapping or human task supervision, achieve competitive results on multiple embodied benchmarks spanning navigation, rearrangement, and manipulation. EnvDreamer also supports image-conditioned reconstruction for real-to-sim studies. Finally, we release EnvDreamer-20k, a dataset of 20,000 validator passed environments with task programs, scene graphs, trajectories, and metadata to support reproducible benchmarking.