Score
Design state representations: define and implement mappings from raw observations and optionally actions into structured, compact encodings or embeddings that preserve task‑relevant information for decision making. This work includes selecting and engineering features, constructing symbolic abstractions and state–action joint encodings, incorporating temporal and semantic features, developing attention‑aware or learned embeddings, and applying compression or coordination‑supporting designs to balance informativeness, efficiency, and downstream learnability.
To address the challenges of learning effective state representations—namely, poor sample efficiency, weak generalization, and difficulty in disentangling semantics—in deep reinforcement learning with complex observation spaces, this paper systematically surveys state-of-the-art model-free online methods. We propose, for the first time, a unified taxonomy comprising six paradigms: self-supervised learning, contrastive learning, causal modeling, information bottleneck, predictive modeling, and disentangled representation learning. Each paradigm is rigorously characterized along six dimensions: objective function, signal source, structural constraints, optimization mechanism, evaluation protocol, and applicable scenarios. The framework explicitly delineates mechanistic principles, strengths, and limitations of existing approaches, and integrates reproducible benchmarks alongside open research challenges. Our analysis significantly enhances representation interpretability, cross-task generalization, and policy robustness. This work establishes the first scalable, theory-grounded analytical guide for state representation learning in deep RL.
Existing latent variable models often suffer from under-constrained objectives, leading to non-identifiable, ambiguous, and poorly interpretable representations. This work proposes the Constrained Latent State Modeling (CLSM) framework, which systematically integrates six core constraints—namely predictive sufficiency, minimality, temporal consistency, and others—for the first time. Grounded in information theory and dynamical systems theory, CLSM formally characterizes the intrinsic couplings and trade-offs among these constraints. By reframing representation learning as a constrained optimization problem, the framework unifies diverse approaches such as variational autoencoders and state-space models, revealing that non-identifiability stems from insufficient constraints rather than technical shortcomings. CLSM thus provides a principled foundation for designing latent variable models that are interpretable, robust, and aligned with downstream tasks.
Current AI agents exhibit weak decision-making generalization in complex environments, primarily due to semantically unstructured and functionally unabstracted state-action representations. To address this, we propose Structurally Enhanced Trajectories (SETs), the first framework to model agent trajectories as multi-layered graph structures encoding object relations, interaction patterns, and functional affordances—thereby overcoming fundamental limitations of sequential modeling. Our approach integrates heterogeneous graph neural networks, hierarchical relational modeling, reinforcement learning–driven trajectory generation, and a novel structured memory architecture, SETLE. Experiments demonstrate that SETs significantly improve structural pattern recognition, semantic interpretability, and function-level transfer across diverse environments. Critically, SETs achieve unified gains in generalization, abstraction, and explainability—advancing all three dimensions simultaneously.
Formal specification of embodied agent behaviors remains challenging due to the difficulty of rigorously encoding perception-driven, dynamic actions within traditional logical frameworks. Method: This paper introduces Embedding Temporal Logic (ETL), the first formal specification language that directly integrates semantic embeddings from pretrained vision and multimodal foundation models. ETL defines behavioral properties as distances between ideal behavior representations and actual observed representations in the embedding space, thereby overcoming limitations of classical logics in capturing perceptually grounded dynamics. The approach unifies embedding-guided specification, formal verification, robot planning, and foundation-model-based control. Contribution/Results: Evaluated on large-model-driven robotic tasks, ETL significantly enhances behavioral interpretability and controllability, enabling precise modeling and targeted steering of desired behaviors while preserving formal guarantees.
Robots struggle to acquire transferable relational concepts from few unannotated, unsegmented demonstrations, resulting in poor zero-shot generalization to long-horizon, high-complexity novel tasks. Method: We propose an end-to-end “real-valued → logical → planning” closed-loop framework that integrates relational learning, unsupervised representation disentanglement, symbolic induction, and differentiable logical reasoning. It autonomously discovers relational symbolic vocabularies, action semantics, and PDDL-like domain models directly from raw continuous trajectories. Contribution: This is the first approach to generate high-level relational concepts fully autonomously—without manual abstraction. We uncover implicit action structures surpassing classical priors. With only a few demonstration trajectories, it constructs high-fidelity abstract models, significantly improving scalability and success rates for joint long-horizon task and motion planning in deterministic environments.
This work investigates the fundamental expressive capacity of linear state-space models (SSMs) for language modeling, clarifying their theoretical modeling boundaries relative to Transformers and classical RNNs. Method: Leveraging formal language theory and automata theory, we formally characterize SSM expressivity—proving for the first time that linear SSMs can exactly recognize star-free languages and optimally model bounded hierarchical structures in memory. We identify a critical expressivity bottleneck in contemporary SSM designs arising from the absence of nonlinearity in state updates. Contribution/Results: Our analysis reveals that SSMs and Transformers possess complementary—not substitutive—capabilities. Empirical evaluation on the Mamba architecture demonstrates substantial gains over Transformers on star-free language tasks and superior memory efficiency in hierarchical structure modeling. These findings provide both theoretical foundations and practical guidance for designing next-generation efficient large language model architectures.
This study formalizes world model research as the problem of designing latent states under task-sufficiency constraints, aiming to retain essential information while discarding redundancy. To this end, it introduces a functional taxonomy of latent states—categorized by roles such as prediction, control, and planning—and develops a seven-dimensional evaluation framework centered on task sufficiency. The framework encompasses diverse technical approaches, including predictive embeddings, recurrent belief states, causal structures, latent action interfaces, embodied planning interfaces, and memory substrates. Empirical results demonstrate that latent state designs aligned with specific task requirements substantially outperform generic modeling strategies that prioritize maximal information preservation, thereby revealing critical distinctions between predictive sufficiency and control sufficiency.
Existing state abstraction methods lack a general principle for rigorously preserving behavioral structure. This work proposes a unified framework that defines behavioral semantics in reinforcement learning in a compositional manner, grounded in local one-step descriptions of system dynamics, and establishes a theory for safe transfer of behavioral structure between abstract and concrete systems. For the first time, the framework enables a compositional formalization of behavioral semantics, supporting the derivation of quantitative metrics with correctness guarantees from logical semantics. It thus lays a principled foundation for behavioral reasoning under state abstraction and provides reusable definitions and provably faithful transfer mechanisms applicable to a broad class of behavioral structures.
This study addresses the challenge of learning effective strategy representations in two-player zero-sum imperfect-information games. To this end, the authors construct a strategy dataset, design a self-supervised embedding method, and evaluate the quality of the learned representations through downstream tasks. Experiments on Kuhn and Leduc poker demonstrate that the proposed approach successfully captures the behavioral and semantic characteristics of strategies. This work is the first to systematically introduce self-supervised learning into the domain of game-theoretic strategy representation, offering a comprehensive framework for strategy embedding learning and evaluation. The results validate the feasibility and potential of this paradigm under fundamental settings.
Existing long-horizon mobile GUI agents struggle to distinguish persistent task states from transient screen observations, often leading to goal forgetting or hallucination. This work proposes a Task State Representation (TSR) framework that explicitly decouples task state from perceptual input through a lightweight external wrapper, without modifying the underlying model architecture or requiring fine-tuning. TSR integrates a global instruction summary, a dynamic subgoal tracker, and an action validator, continuously updating the task state via visual comparisons before and after each action. This approach achieves, for the first time, training-agnostic separation of task state and observation, significantly boosting performance across four mobile GUI benchmarks—improving success rates by up to 12 percentage points on complex cross-app and memory-intensive tasks.