JEPA for RL: Investigating Joint-Embedding Predictive Architectures for Reinforcement Learning

๐Ÿ“… 2025-04-23
๐Ÿ›๏ธ ESANN 2025 proceesdings
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the representation collapse problem in vision-based reinforcement learning (RL), caused by the entanglement of visual representation learning and policy optimization. We introduce the Joint-Embedding Predictive Architecture (JEPA)โ€”a self-supervised frameworkโ€”into RL for the first time, proposing a decoupled representation learning mechanism: a vision Transformer is employed to construct JEPAโ€™s predictive objective, explicitly separating perceptual modeling from policy optimization to mitigate representation degradation. Evaluated on dynamic control benchmarks including CartPole, our approach significantly improves training stability. The robust, JEPA-derived visual embeddings serve as high-quality inputs for downstream policy learning, enabling end-to-end policies with superior performance and generalization. This work establishes a novel paradigm for self-supervised, representation-driven visual RL.

Technology Category

Computer Vision: Representation Learning for VisionMachine Learning: Representation LearningIntelligent Robots: Learning & Optimization for ROB

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsResponsible Web: Machine-in-the-loop, human agency and autonomy
๐Ÿ“ Abstract
Joint-Embedding Predictive Architectures (JEPA) have recently become popular as promising architectures for self-supervised learning. Vision transformers have been trained using JEPA to produce embeddings from images and videos, which have been shown to be highly suitable for downstream tasks like classification and segmentation. In this paper, we show how to adapt the JEPA architecture to reinforcement learning from images. We discuss model collapse, show how to prevent it, and provide exemplary data on the classical Cart Pole task.
Problem

Research questions and friction points this paper is trying to address.

Adapt JEPA architecture for reinforcement learning from images
Address model collapse in JEPA-based reinforcement learning
Demonstrate effectiveness on classical Cart Pole task
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adapt JEPA to reinforcement learning from images
Prevent model collapse in JEPA architecture
Apply JEPA to classical Cart Pole task
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
T
Tristan Kenneweg
University of Bielefeld - Technical Faculty
P
Philip Kenneweg
University of Bielefeld - Technical Faculty
Barbara Hammer
Barbara Hammer
Professor, Bielefeld University
machine learningdata miningneural networksbioinformaticstheoretical computer science