SimVLA: Zero-Shot Sim-to-Real VLA Learning for Mobile Manipulation

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of Vision-Language-Action (VLA) models on expensive real-world data and the underexplored potential of simulation by proposing an end-to-end, simulation-only training framework that achieves zero-shot sim-to-real transfer for mobile manipulation. Methodologically, it introduces a novel training paradigm combining atomic skill generation with privileged state supervision. Through multi-source simulation pre-training and policy rollback post-training, the approach integrates spatial geometry with subtask-level vision-language supervision, entirely eliminating dependence on teleoperation. Experimental results demonstrate that this framework outperforms policies trained on 50 real-world demonstrations in complex tasks such as household restocking, strongly validating the scalability of synthetic data-driven VLA models.
📝 Abstract
Large-scale, diverse datasets have driven the success of LLMs and VLMs. But VLAs for robotics remain limited by the cost and complexity of real-world data collection. While simulation offers a scalable alternative, its potential for sim-to-real VLA learning in mobile manipulation remains largely underexplored. We introduce SimVLA, an end-to-end framework that trains VLAs entirely on synthetic simulation data without teleoperation for mobile manipulation. SimVLA is first pre-trained on two complementary simulation-derived datasets: SimAction, a large-scale robot action dataset spanning 35 diverse mobile manipulation tasks, generated by composing atomic skills, and SimVQA, which leverages privileged simulator state to provide spatial, geometric, and subtask-level visual-language supervision. We further post-train SimVLA on a mixture of SimAction and SimDeploy, a dataset collected from policy rollouts across diverse simulated environments. We evaluate SimVLA on tasks including restocking, pouring, and cleaning, and show zero-shot transfer to real-world mobile manipulation, including real home environments. SimVLA outperforms policies trained on 50 in-domain real-world demonstrations, suggesting that simulation can enable scalable sim-to-real mobile manipulation. We further demonstrate the value of multiple complementary forms of supervision for effectively leveraging simulation in VLA training.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
Sim-to-Real
Mobile Manipulation
Zero-Shot Transfer
Synthetic Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action (VLA)
Sim-to-Real Transfer
Zero-Shot Learning
Mobile Manipulation
Synthetic Data
🔎 Similar Papers
No similar papers found.