Score
Design, build, and evaluate encoder architectures and self‑supervised objectives that produce compact, informative embeddings from unlabeled data, covering visual, video, set-structured, discrete, and temporally irregular inputs. These methods include predictive and auxiliary tasks (including for reinforcement learning), nonlinear mapping strategies to handle missing or noisy samplings, and evaluation pipelines to ensure the representations are useful for downstream prediction, control, transfer, and other tasks.
This work addresses the representation collapse problem in vision-based reinforcement learning (RL), caused by the entanglement of visual representation learning and policy optimization. We introduce the Joint-Embedding Predictive Architecture (JEPA)—a self-supervised framework—into RL for the first time, proposing a decoupled representation learning mechanism: a vision Transformer is employed to construct JEPA’s predictive objective, explicitly separating perceptual modeling from policy optimization to mitigate representation degradation. Evaluated on dynamic control benchmarks including CartPole, our approach significantly improves training stability. The robust, JEPA-derived visual embeddings serve as high-quality inputs for downstream policy learning, enabling end-to-end policies with superior performance and generalization. This work establishes a novel paradigm for self-supervised, representation-driven visual RL.
This work addresses the challenges of state representation and poor sample efficiency in visual reinforcement learning caused by high-dimensional image inputs. The authors propose a self-supervised auxiliary task based on masked prediction, which leverages observation sequences and their contextual information collected by the agent to learn compact yet informative sequential representations in a latent space using a Transformer architecture. Unlike conventional reconstruction-based approaches, this method employs a non-reconstructive masked prediction mechanism that emphasizes understanding of task dynamics rather than pixel-level fidelity. Experimental results demonstrate that the proposed approach significantly outperforms state-of-the-art sample-efficient reinforcement learning algorithms across multiple continuous and discrete control benchmarks, achieving substantial gains in sample efficiency.
This work addresses the limitations of traditional generative modeling, which often focuses on pixel-level reconstruction and struggles to capture high-level semantics. To overcome this, the authors propose an Energy-Based Joint Embedding Predictive Architecture (EB-JEPA) that performs self-supervised prediction in representation space rather than pixel space, effectively enabling the construction of world models for images, videos, and action-conditioned environments. The study introduces the first open-source, lightweight, and modular EB-JEPA library, systematically demonstrating the critical role of regularization in preventing representational collapse. The framework supports multi-step temporal prediction and action-conditioned modeling, achieving strong empirical results: 91% probe accuracy on CIFAR-10, high-quality multi-step video prediction on Moving MNIST, and a 97% planning success rate on the Two Rooms navigation task—all trained within hours on a single GPU.
To address the instability of adversarial training and bias inherent in handcrafted reward functions in unsupervised video summarization, this paper proposes a two-stage decoupled reinforcement learning framework. In the first stage, a self-supervised pre-trained video reconstruction model generates frame-level reconstruction fidelity scores, serving as a learnable, data-driven reward signal. In the second stage, Proximal Policy Optimization (PPO) optimizes an importance-weighted summarization policy, with end-to-end differentiable approximation enabling efficient training. Crucially, the method eliminates heuristic reward design and adversarial training, being the first to directly model reconstruction quality as the RL reward. Evaluated on TVSum and SumMe, it achieves F-scores of 62.3 and 54.5, respectively—outperforming prior work—while accelerating inference by 300× over the state-of-the-art. Moreover, the generated summary distributions exhibit significantly improved alignment with human annotations.
This work addresses reward-free, offline, image-driven goal-oriented robotic manipulation—enabling end-to-end vision-based grasping and relocation of real-world objects using only a single target image. Methodologically, we propose a contrastive learning–based self-supervised offline reinforcement learning framework, integrating an image encoder–action decoder architecture with goal-conditioned policy learning; crucial architectural designs and hyperparameter configurations are introduced to ensure stable training for real-hardware deployment. To the best of our knowledge, this is the first demonstration of contrastive self-supervised RL on a physical robotic arm. Our approach achieves a twofold improvement in task success rate over baseline methods. Critically, it requires no handcrafted reward functions, online environment interaction, or pixel-level annotations—significantly lowering the barrier to real-world deployment.
This study addresses the challenge of learning compact and semantically meaningful representations of chess positions from continuous game sequences under unsupervised conditions. To this end, the authors propose a novel self-supervised learning framework that integrates concepts from Masked Autoencoders (MAE), Joint-Embedding Predictive Architecture (JEPA), and BERT. By predicting masked board states within a low-dimensional embedding space, the model effectively encodes positional semantics without relying on reinforcement learning or explicit move labels. This work represents the first application of a combined MAE–JEPA–BERT architecture to sequential board-game modeling, enabling the capture of piece movement logic purely through self-supervision. Experimental results demonstrate that the learned representation space naturally clusters into interpretable, chess-theoretic concepts, clearly reflects positional semantics, and exhibits the capacity to reason about legal moves.
While contemporary self-supervised and masked/denoising autoencoder methods effectively learn strong representations from massive unlabeled data, their representational nature, cross-task generalization capability, and emergence mechanisms remain theoretically unexplained. Method: This project integrates statistical inference and nonconvex optimization theory to establish a unified analytical framework for unsupervised representation learning. Contribution/Results: It provides the first mathematical characterization of how self-supervised objectives—such as contrastive learning and reconstruction losses—induce structured latent spaces, and quantitatively links representation linear separability, invariance, and downstream generalization. The work identifies key theoretical conditions under which pretrained models achieve zero-shot transfer and task emergence in vision foundation models. Crucially, it delivers the first theoretical foundation for large-scale pretraining that is both statistically interpretable and optimization-traceable—bridging statistical guarantees with practical training dynamics.
This work proposes a novel self-supervised visual representation learning paradigm, Temporal Difference in Vision (TDV), which eschews strong inductive biases such as data augmentation, masking, or cropping commonly used in existing methods. Instead, TDV leverages the temporal causal assumption that “the past causes the future” in videos, jointly training an image encoder and a motion encoder so that the sum of the current frame’s representation and the motion representation approximates the representation of the subsequent frame. Relying solely on this weak temporal assumption, the method achieves state-of-the-art performance among self-supervised approaches on dense spatial tasks, demonstrating the effectiveness and potential of modeling temporal causality for large-scale visual representation learning.
This work addresses the lack of a unified theoretical foundation in unsupervised visual representation learning, where existing methods struggle to simultaneously achieve semantic invariance, spatial structure modeling, and non-degenerate solutions. The authors propose three essential principles—observation, prediction, and regularization—and formulate them within a unified energy-based decomposition framework, offering the first formalization of core self-supervised learning criteria. Through rigorous analysis of gradient complementarity, convergence guarantees for momentum encoders, and a negative-sample-free alignment theory, the study exposes fundamental limitations of contrastive learning and momentum mechanisms, demonstrating that all three principles are indispensable. Controlled experiments, including block retrieval evaluations, confirm that optimal performance is attained only when these principles operate in concert.
This work addresses the lack of systematic investigation into extremely compact neural video representations, as existing approaches predominantly focus on medium- to high-capacity models. The study presents TinyNeRV, a novel architecture that establishes the performance limits of minimal-scale NeRVs through integrated strategies including capacity scaling, frequency-aware knowledge distillation, and low-precision inference—encompassing both post-training quantization and quantization-aware training. These techniques collectively achieve substantial reductions in model parameters, computational cost, and memory footprint. Extensive experiments across multiple video datasets demonstrate that TinyNeRV attains an exceptional trade-off between reconstruction quality and efficiency, thereby validating the feasibility and robustness of lightweight neural video representations in resource-constrained and real-time deployment scenarios.