π€ AI Summary
This study addresses the low sample efficiency and limited control capabilities of existing Vision-Language-Action (VLA) models, which rely on massive datasets and lack explicit world models. We propose training a predictive world model within the embedding space of the VLAβs visual encoder. Drawing upon the Joint Embedding Predictive Architecture (JEPA), our approach discards pixel-level reconstruction losses in favor of predicting future states via action-conditioned embeddings. This method demonstrates that visual embeddings provide a non-lossy representation for dynamics simulation, significantly reducing data requirements while enhancing sample efficiency. Furthermore, it enables short-horizon planning during inference by generating goal images to guide VLA action outputs, thereby fully exploiting the representational capacity of visual embeddings.
π Abstract
Vision-Language-Action models (VLAs), while showing strong potential for robot control, require massive amounts of high-quality imitation learning data. Moreover, the absence of an explicit world model casts further doubt on their control capabilities. In this concept paper, we propose a novel architecture that addresses sample efficiency in VLAs by training a predictive world model on the embedding space of the VLA's vision encoder. We hypothesize that these embeddings are action-relevant and usable for future prediction. To this end, we propose using the suggested architecture to investigate how well these embeddings predict the future based on actions, as the inability to do so would mark a key limitation of VLA architectures: the lack of a non-lossy implicit world model to simulate real-world dynamics. The proposed architecture differs from the standard world model dynamics as the loss comes from the embedding space rather than the pixel space, similar to joint embedding predictive architectures. Furthermore, the trained world model can be utilized for short-term planning tasks by sampling VLA actions given goal images. We intend to examine the richness of vision embeddings in VLAs and reduce their high data requirements through a world model that can also generate plans during inference.