self-supervised representation learning

Design, build, and evaluate encoder architectures and self‑supervised objectives that produce compact, informative embeddings from unlabeled data, covering visual, video, set-structured, discrete, and temporally irregular inputs. These methods include predictive and auxiliary tasks (including for reinforcement learning), nonlinear mapping strategies to handle missing or noisy samplings, and evaluation pipelines to ensure the representations are useful for downstream prediction, control, transfer, and other tasks.

self-supervisedrepresentationlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the representation collapse problem in vision-based reinforcement learning (RL), caused by the entanglement of visual representation learning and policy optimization. We introduce the Joint-Embedding Predictive Architecture (JEPA)—a self-supervised framework—into RL for the first time, proposing a decoupled representation learning mechanism: a vision Transformer is employed to construct JEPA’s predictive objective, explicitly separating perceptual modeling from policy optimization to mitigate representation degradation. Evaluated on dynamic control benchmarks including CartPole, our approach significantly improves training stability. The robust, JEPA-derived visual embeddings serve as high-quality inputs for downstream policy learning, enabling end-to-end policies with superior performance and generalization. This work establishes a novel paradigm for self-supervised, representation-driven visual RL.

Adapt JEPA architecture for reinforcement learning from imagesAddress model collapse in JEPA-based reinforcement learningDemonstrate effectiveness on classical Cart Pole task

This work addresses the challenges of state representation and poor sample efficiency in visual reinforcement learning caused by high-dimensional image inputs. The authors propose a self-supervised auxiliary task based on masked prediction, which leverages observation sequences and their contextual information collected by the agent to learn compact yet informative sequential representations in a latent space using a Transformer architecture. Unlike conventional reconstruction-based approaches, this method employs a non-reconstructive masked prediction mechanism that emphasizes understanding of task dynamics rather than pixel-level fidelity. Experimental results demonstrate that the proposed approach significantly outperforms state-of-the-art sample-efficient reinforcement learning algorithms across multiple continuous and discrete control benchmarks, achieving substantial gains in sample efficiency.

high-dimensional inputssample efficiencyself-supervised learning

This work addresses the limitations of traditional generative modeling, which often focuses on pixel-level reconstruction and struggles to capture high-level semantics. To overcome this, the authors propose an Energy-Based Joint Embedding Predictive Architecture (EB-JEPA) that performs self-supervised prediction in representation space rather than pixel space, effectively enabling the construction of world models for images, videos, and action-conditioned environments. The study introduces the first open-source, lightweight, and modular EB-JEPA library, systematically demonstrating the critical role of regularization in preventing representational collapse. The framework supports multi-step temporal prediction and action-conditioned modeling, achieving strong empirical results: 91% probe accuracy on CIFAR-10, high-quality multi-step video prediction on Moving MNIST, and a 97% planning success rate on the Two Rooms navigation task—all trained within hours on a single GPU.

Energy-Based ModelsJoint-Embedding Predictive ArchitecturesRepresentation Learning

To address the instability of adversarial training and bias inherent in handcrafted reward functions in unsupervised video summarization, this paper proposes a two-stage decoupled reinforcement learning framework. In the first stage, a self-supervised pre-trained video reconstruction model generates frame-level reconstruction fidelity scores, serving as a learnable, data-driven reward signal. In the second stage, Proximal Policy Optimization (PPO) optimizes an importance-weighted summarization policy, with end-to-end differentiable approximation enabling efficient training. Crucially, the method eliminates heuristic reward design and adversarial training, being the first to directly model reconstruction quality as the RL reward. Evaluated on TVSum and SumMe, it achieves F-scores of 62.3 and 54.5, respectively—outperforming prior work—while accelerating inference by 300× over the state-of-the-art. Moreover, the generated summary distributions exhibit significantly improved alignment with human annotations.

Addresses unstable adversarial training and heuristic reward functionsUnsupervised video summarization using reinforcement learningUses reconstruction fidelity as proxy for summary informativeness

Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data

Jun 06, 2023
CZ
Chongyi Zheng
🏛️ Carnegie Mellon University | Princeton University | UC Berkeley | University of Washington | Cornell University

This work addresses reward-free, offline, image-driven goal-oriented robotic manipulation—enabling end-to-end vision-based grasping and relocation of real-world objects using only a single target image. Methodologically, we propose a contrastive learning–based self-supervised offline reinforcement learning framework, integrating an image encoder–action decoder architecture with goal-conditioned policy learning; crucial architectural designs and hyperparameter configurations are introduced to ensure stable training for real-hardware deployment. To the best of our knowledge, this is the first demonstration of contrastive self-supervised RL on a physical robotic arm. Our approach achieves a twofold improvement in task success rate over baseline methods. Critically, it requires no handcrafted reward functions, online environment interaction, or pixel-level annotations—significantly lowering the barrier to real-world deployment.

Enabling real-world image-based robotic manipulation without human labelsImproving success rates via architecture and hyperparameter tuningStabilizing self-supervised RL for robotic goal reaching

Latest Papers

What's happening recently
View more

This study addresses the challenge of learning compact and semantically meaningful representations of chess positions from continuous game sequences under unsupervised conditions. To this end, the authors propose a novel self-supervised learning framework that integrates concepts from Masked Autoencoders (MAE), Joint-Embedding Predictive Architecture (JEPA), and BERT. By predicting masked board states within a low-dimensional embedding space, the model effectively encodes positional semantics without relying on reinforcement learning or explicit move labels. This work represents the first application of a combined MAE–JEPA–BERT architecture to sequential board-game modeling, enabling the capture of piece movement logic purely through self-supervision. Experimental results demonstrate that the learned representation space naturally clusters into interpretable, chess-theoretic concepts, clearly reflects positional semantics, and exhibits the capacity to reason about legal moves.

chesslatent spacerepresentation learning

Theoretical Foundations of Representation Learning using Unlabeled Data: Statistics and Optimization

Sep 23, 2025
PE
Pascal Esser
🏛️ Ludwig-Maximilians-Universität München | Technical University of Munich

While contemporary self-supervised and masked/denoising autoencoder methods effectively learn strong representations from massive unlabeled data, their representational nature, cross-task generalization capability, and emergence mechanisms remain theoretically unexplained. Method: This project integrates statistical inference and nonconvex optimization theory to establish a unified analytical framework for unsupervised representation learning. Contribution/Results: It provides the first mathematical characterization of how self-supervised objectives—such as contrastive learning and reconstruction losses—induce structured latent spaces, and quantitatively links representation linear separability, invariance, and downstream generalization. The work identifies key theoretical conditions under which pretrained models achieve zero-shot transfer and task emergence in vision foundation models. Crucially, it delivers the first theoretical foundation for large-scale pretraining that is both statistically interpretable and optimization-traceable—bridging statistical guarantees with practical training dynamics.

Analyzing deep unsupervised representation learning principles using classical theoriesCharacterizing representations learned by self-supervision and masked autoencodersExplaining why these models perform well across diverse prediction tasks

This work proposes a novel self-supervised visual representation learning paradigm, Temporal Difference in Vision (TDV), which eschews strong inductive biases such as data augmentation, masking, or cropping commonly used in existing methods. Instead, TDV leverages the temporal causal assumption that “the past causes the future” in videos, jointly training an image encoder and a motion encoder so that the sum of the current frame’s representation and the motion representation approximates the representation of the subsequent frame. Relying solely on this weak temporal assumption, the method achieves state-of-the-art performance among self-supervised approaches on dense spatial tasks, demonstrating the effectiveness and potential of modeling temporal causality for large-scale visual representation learning.

Inductive BiasesSelf-Supervised LearningTemporal Differences

This work addresses the lack of a unified theoretical foundation in unsupervised visual representation learning, where existing methods struggle to simultaneously achieve semantic invariance, spatial structure modeling, and non-degenerate solutions. The authors propose three essential principles—observation, prediction, and regularization—and formulate them within a unified energy-based decomposition framework, offering the first formalization of core self-supervised learning criteria. Through rigorous analysis of gradient complementarity, convergence guarantees for momentum encoders, and a negative-sample-free alignment theory, the study exposes fundamental limitations of contrastive learning and momentum mechanisms, demonstrating that all three principles are indispensable. Controlled experiments, including block retrieval evaluations, confirm that optimal performance is attained only when these principles operate in concert.

non-degeneracyself-supervised learningsemantic invariance

This work addresses the lack of systematic investigation into extremely compact neural video representations, as existing approaches predominantly focus on medium- to high-capacity models. The study presents TinyNeRV, a novel architecture that establishes the performance limits of minimal-scale NeRVs through integrated strategies including capacity scaling, frequency-aware knowledge distillation, and low-precision inference—encompassing both post-training quantization and quantization-aware training. These techniques collectively achieve substantial reductions in model parameters, computational cost, and memory footprint. Extensive experiments across multiple video datasets demonstrate that TinyNeRV attains an exceptional trade-off between reconstruction quality and efficiency, thereby validating the feasibility and robustness of lightweight neural video representations in resource-constrained and real-time deployment scenarios.

compact modelsmodel capacity scalingneural video representations

Hot Scholars

HJ

Han-Jia Ye

Nanjing University
Machine LearningData MiningMetric LearningMeta-Learning
SZ

Shizhen Zhao

Associate Professor, John Hopcroft Center, Shanghai Jiao Tong Univerisity
Hybrid Electrical/Optical Data Center NetworksDeterministic NetworksNetwork Optimization
TT

Tao Tan

FCA MPU
Medical Imaging AI
YM

Yuki M. Asano

Full Professor, Head of FunAI Lab, University of Technology Nuremberg
Deep LearningMultimodal LearningSelf-supervised LearningLarge Model Adaptation
TT

Tinne Tuytelaars

KU Leuven - PSI, Belgium
computer visioncontinual learning