๐ค AI Summary
This work addresses the representation collapse problem in vision-based reinforcement learning (RL), caused by the entanglement of visual representation learning and policy optimization. We introduce the Joint-Embedding Predictive Architecture (JEPA)โa self-supervised frameworkโinto RL for the first time, proposing a decoupled representation learning mechanism: a vision Transformer is employed to construct JEPAโs predictive objective, explicitly separating perceptual modeling from policy optimization to mitigate representation degradation. Evaluated on dynamic control benchmarks including CartPole, our approach significantly improves training stability. The robust, JEPA-derived visual embeddings serve as high-quality inputs for downstream policy learning, enabling end-to-end policies with superior performance and generalization. This work establishes a novel paradigm for self-supervised, representation-driven visual RL.
๐ Abstract
Joint-Embedding Predictive Architectures (JEPA) have recently become popular as promising architectures for self-supervised learning. Vision transformers have been trained using JEPA to produce embeddings from images and videos, which have been shown to be highly suitable for downstream tasks like classification and segmentation. In this paper, we show how to adapt the JEPA architecture to reinforcement learning from images. We discuss model collapse, show how to prevent it, and provide exemplary data on the classical Cart Pole task.