🤖 AI Summary
This study addresses the reliance of Vision-Language-Action (VLA) models on expensive real-world data and the underexplored potential of simulation by proposing an end-to-end, simulation-only training framework that achieves zero-shot sim-to-real transfer for mobile manipulation. Methodologically, it introduces a novel training paradigm combining atomic skill generation with privileged state supervision. Through multi-source simulation pre-training and policy rollback post-training, the approach integrates spatial geometry with subtask-level vision-language supervision, entirely eliminating dependence on teleoperation. Experimental results demonstrate that this framework outperforms policies trained on 50 real-world demonstrations in complex tasks such as household restocking, strongly validating the scalability of synthetic data-driven VLA models.
📝 Abstract
Large-scale, diverse datasets have driven the success of LLMs and VLMs. But VLAs for robotics remain limited by the cost and complexity of real-world data collection. While simulation offers a scalable alternative, its potential for sim-to-real VLA learning in mobile manipulation remains largely underexplored. We introduce SimVLA, an end-to-end framework that trains VLAs entirely on synthetic simulation data without teleoperation for mobile manipulation. SimVLA is first pre-trained on two complementary simulation-derived datasets: SimAction, a large-scale robot action dataset spanning 35 diverse mobile manipulation tasks, generated by composing atomic skills, and SimVQA, which leverages privileged simulator state to provide spatial, geometric, and subtask-level visual-language supervision. We further post-train SimVLA on a mixture of SimAction and SimDeploy, a dataset collected from policy rollouts across diverse simulated environments. We evaluate SimVLA on tasks including restocking, pouring, and cleaning, and show zero-shot transfer to real-world mobile manipulation, including real home environments. SimVLA outperforms policies trained on 50 in-domain real-world demonstrations, suggesting that simulation can enable scalable sim-to-real mobile manipulation. We further demonstrate the value of multiple complementary forms of supervision for effectively leveraging simulation in VLA training.