Score
Designs and trains models that encode sequences of observations and actions into compact belief-state representations (e.g., latent vectors or distributions) which capture uncertainty and the sufficient statistics needed for downstream decision-making. This includes developing and evaluating pretrained belief encoders, representation-learning objectives, and interfaces that decouple representation learning from policy learning so policies consume and act on learned belief states.
To address the challenges of learning effective state representations—namely, poor sample efficiency, weak generalization, and difficulty in disentangling semantics—in deep reinforcement learning with complex observation spaces, this paper systematically surveys state-of-the-art model-free online methods. We propose, for the first time, a unified taxonomy comprising six paradigms: self-supervised learning, contrastive learning, causal modeling, information bottleneck, predictive modeling, and disentangled representation learning. Each paradigm is rigorously characterized along six dimensions: objective function, signal source, structural constraints, optimization mechanism, evaluation protocol, and applicable scenarios. The framework explicitly delineates mechanistic principles, strengths, and limitations of existing approaches, and integrates reproducible benchmarks alongside open research challenges. Our analysis significantly enhances representation interpretability, cross-task generalization, and policy robustness. This work establishes the first scalable, theory-grounded analytical guide for state representation learning in deep RL.
Existing adaptive data acquisition methods often suffer from inefficient policy learning due to reliance on biased posterior approximations or inadequate exploitation of model representations. This work proposes POLAR, a novel framework that decouples representation learning from policy learning by leveraging pretrained predictive foundation models to encode belief states. POLAR unifies Bayesian experimental design, Bayesian optimization, and active learning within a single coherent framework. By integrating amortized policy learning with task-specific utility functions, the approach substantially reduces the required number of training samples and consistently outperforms state-of-the-art amortized methods across diverse tasks, significantly enhancing the scalability and efficiency of data acquisition.
In partially observable meta-reinforcement learning (meta-RL), existing approaches approximate Bayesian-optimal policies but struggle to learn compact, interpretable belief-state representations, limiting generalization and adaptation. This work introduces predictive coding—a neuroscientific principle—into meta-RL for the first time, proposing a self-supervised predictive coding module that jointly optimizes history compression and auxiliary prediction objectives to enable efficient representation learning over observation histories in state-machine environments. The method preserves policy performance while substantially improving both interpretability and Bayesian optimality of belief representations. On benchmark tasks such as active information seeking, it is the only approach achieving simultaneously optimal policies and optimal representations. Moreover, it demonstrates significantly enhanced cross-task generalization.
To address insufficient uncertainty modeling in state representations for partially observable reinforcement learning (RL), this paper introduces the Differentiable Kalman Filter (DKF) layer—a plug-and-play state-space module. The DKF layer explicitly models latent states as Gaussian distributions, enabling closed-form probabilistic inference, and achieves efficient differentiability via a parallel scan algorithm, allowing seamless integration into model-agnostic, end-to-end RL frameworks. Unlike conventional RNNs or Transformers—which lack explicit probabilistic filtering mechanisms—the DKF layer is the first to reformulate analytical Kalman filtering as a learnable, parallelizable state representation unit. Experiments across diverse partially observable benchmarks demonstrate that DKF significantly outperforms LSTM, GRU, and Transformer baselines; notably, on tasks demanding deep uncertainty reasoning, it yields 17–32% higher cumulative returns.
This study addresses the challenge of learning effective strategy representations in two-player zero-sum imperfect-information games. To this end, the authors construct a strategy dataset, design a self-supervised embedding method, and evaluate the quality of the learned representations through downstream tasks. Experiments on Kuhn and Leduc poker demonstrate that the proposed approach successfully captures the behavioral and semantic characteristics of strategies. This work is the first to systematically introduce self-supervised learning into the domain of game-theoretic strategy representation, offering a comprehensive framework for strategy embedding learning and evaluation. The results validate the feasibility and potential of this paradigm under fundamental settings.
This work addresses probabilistic inference—forecasting, abduction, and intermediate-state estimation—for high-dimensional time series. We propose a probabilistic representation framework grounded in temporal contrastive learning. We provide the first theoretical proof that the learned latent states follow a Gaussian Markov chain structure, thereby reducing complex probabilistic inference to closed-form algebraic operations in a low-dimensional space: matrix inversion for abduction and linear interpolation for intermediate-state estimation. Integrating temporal contrastive learning, Gaussian graphical models, and linear-algebraic inference, our approach enables analytically tractable, efficient, and interpretable probabilistic reasoning. Empirical validation on synthetic tasks with up to 46 dimensions confirms the validity and scalability of the closed-form solutions, achieving substantial reductions in computational complexity compared to conventional sampling- or optimization-based methods.
Existing latent variable models often suffer from under-constrained objectives, leading to non-identifiable, ambiguous, and poorly interpretable representations. This work proposes the Constrained Latent State Modeling (CLSM) framework, which systematically integrates six core constraints—namely predictive sufficiency, minimality, temporal consistency, and others—for the first time. Grounded in information theory and dynamical systems theory, CLSM formally characterizes the intrinsic couplings and trade-offs among these constraints. By reframing representation learning as a constrained optimization problem, the framework unifies diverse approaches such as variational autoencoders and state-space models, revealing that non-identifiability stems from insufficient constraints rather than technical shortcomings. CLSM thus provides a principled foundation for designing latent variable models that are interpretable, robust, and aligned with downstream tasks.
This work addresses the challenge of constructing reliable and compact belief representations that support near-optimal decision-making under perceptual and actuation noise. It introduces a “reliability cell covering” approach that replaces traditional equivalence-class partitioning by defining cells in belief space wherein the optimal action-value function varies by no more than a tolerance ε. The method employs a reliability entropy measure to quantify decision-relevant belief complexity and distinguishes between representational sufficiency and performance limits imposed by noise. By leveraging a fixed observation filtering map, predictive observation laws, and a controlled belief transition kernel, the construction yields an ε-cover under Lipschitz continuity assumptions, applicable to finite POMDPs, linear-Gaussian systems, and particle-filter-based models. The resulting piecewise-constant policies achieve a suboptimality bound of 2ε/(1−γ) and admit analytical or empirical verification across diverse filtering frameworks.
This work addresses a key limitation in existing hierarchical offline goal-conditioned reinforcement learning, where reliance on value-based representations often fails to distinguish states critical for action selection, thereby constraining control performance. The authors propose an information-theoretic framework that explicitly differentiates between “value sufficiency” and “action sufficiency,” demonstrating that the latter is essential for preserving sufficient information between high-level planning and low-level execution to enable optimal action selection. Through information-theoretic analysis, a hierarchical policy architecture, and log-loss–based training of the low-level policy, the resulting action-sufficient representations consistently outperform conventional value-based goal representations across discrete environments and standard benchmarks, empirically validating their strong correlation with improved control performance.
This work addresses the challenge in offline goal-conditioned reinforcement learning where sparse rewards often cause misalignment between state and goal representations, leading encoders to collapse into goal-irrelevant low-dimensional subspaces and degrading policy stability. To mitigate this, the authors propose Ms.PR, a multi-scale representation learning framework that, for the first time, enforces cross-scale predictive consistency as a core constraint to achieve hierarchical alignment in latent space—from local dynamics to long-horizon goal structures. Integrating multi-scale predictive modeling, latent-space alignment constraints, and an offline RL architecture, Ms.PR supports both visual and state-based inputs and consistently enhances representation quality and policy robustness across diverse tasks, trajectory stitching scenarios, and high-noise conditions, outperforming existing methods.
This work addresses the limited generalization of reinforcement learning agents to new tasks, which often necessitates training from scratch. To overcome this, the authors propose Outcome-Predictive State Representations (OPSRs) and the OPSR Skill framework, which construct compact, task-agnostic state abstractions and define reusable abstract actions—referred to as skills—on top of these representations. This approach is the first to jointly abstract both states and actions, enabling cross-task skill transfer without requiring task-specific preprocessing, while preserving policy optimality. Empirical results demonstrate that OPSR-based skills significantly accelerate learning across multiple unseen tasks, confirming their strong generalization capability and effectiveness.