Towards VLA-Dreamer: Refining VLA Behavior Using World Models

πŸ“… 2026-09-25
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the low sample efficiency and limited control capabilities of existing Vision-Language-Action (VLA) models, which rely on massive datasets and lack explicit world models. We propose training a predictive world model within the embedding space of the VLA’s visual encoder. Drawing upon the Joint Embedding Predictive Architecture (JEPA), our approach discards pixel-level reconstruction losses in favor of predicting future states via action-conditioned embeddings. This method demonstrates that visual embeddings provide a non-lossy representation for dynamics simulation, significantly reducing data requirements while enhancing sample efficiency. Furthermore, it enables short-horizon planning during inference by generating goal images to guide VLA action outputs, thereby fully exploiting the representational capacity of visual embeddings.
πŸ“ Abstract
Vision-Language-Action models (VLAs), while showing strong potential for robot control, require massive amounts of high-quality imitation learning data. Moreover, the absence of an explicit world model casts further doubt on their control capabilities. In this concept paper, we propose a novel architecture that addresses sample efficiency in VLAs by training a predictive world model on the embedding space of the VLA's vision encoder. We hypothesize that these embeddings are action-relevant and usable for future prediction. To this end, we propose using the suggested architecture to investigate how well these embeddings predict the future based on actions, as the inability to do so would mark a key limitation of VLA architectures: the lack of a non-lossy implicit world model to simulate real-world dynamics. The proposed architecture differs from the standard world model dynamics as the loss comes from the embedding space rather than the pixel space, similar to joint embedding predictive architectures. Furthermore, the trained world model can be utilized for short-term planning tasks by sampling VLA actions given goal images. We intend to examine the richness of vision embeddings in VLAs and reduce their high data requirements through a world model that can also generate plans during inference.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
world model
sample efficiency
robot control
future prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action models
World Models
Embedding Space Prediction
Short-term Planning
Sample Efficiency
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
P
Parsa Mastouri Kashani
Knowledge Technology, Dept. of Informatics, University of Hamburg, Germany
Jan-Gerrit Habekost
Jan-Gerrit Habekost
University of Hamburg
Neurorobotics
S
Stefan Wermter
Knowledge Technology, Dept. of Informatics, University of Hamburg, Germany