entity-centric rl

Designs and implements reinforcement-learning state representations, architectures, and policies that explicitly represent environment entities or objects; this includes building perception or preprocessing pipelines to disaggregate observations into entity-level feature vectors and encoding relational/topological and temporal configuration information separately. Uses those entity-centric representations to train and evaluate permutation‑invariant or set-based policies and value functions that act over collections of object/entity representations.

entity-centricrl

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Current AI agents exhibit weak decision-making generalization in complex environments, primarily due to semantically unstructured and functionally unabstracted state-action representations. To address this, we propose Structurally Enhanced Trajectories (SETs), the first framework to model agent trajectories as multi-layered graph structures encoding object relations, interaction patterns, and functional affordances—thereby overcoming fundamental limitations of sequential modeling. Our approach integrates heterogeneous graph neural networks, hierarchical relational modeling, reinforcement learning–driven trajectory generation, and a novel structured memory architecture, SETLE. Experiments demonstrate that SETs significantly improve structural pattern recognition, semantic interpretability, and function-level transfer across diverse environments. Critically, SETs achieve unified gains in generalization, abstraction, and explainability—advancing all three dimensions simultaneously.

Developing structurally enriched trajectory representationsEnhancing AI decision-making in complex environmentsImproving generalization across diverse tasks and domains

Relational Object-Centric Actor-Critic

Oct 26, 2023
LU
L. Ugadiarov
🏛️ FRC CSC RAS | MIPT | AIRI

Existing object-centric model-free and monolithic model-based methods exhibit limited generalization in image-driven reinforcement learning tasks involving numerous objects and complex dynamics. Method: We propose a novel algorithm integrating an object-centric world model with the Actor-Critic framework. Our approach jointly optimizes the world model and policy by unifying object-centric representation learning, relational inductive neural networks, and model-based RL. Crucially, it embeds predictive state/reward modeling directly into the Critic network and explicitly models actions as causal interventions on object relations—formulating model learning as causal relation induction. Contribution/Results: This is the first work to incorporate causal intervention modeling and predictive Critic design within an object-centric model-based RL framework. Experiments in 3D robotic manipulation and 2D compositional structure environments demonstrate substantial improvements over state-of-the-art object-centric model-free and monolithic model-based baselines—achieving up to 37% performance gain under high object density and intricate dynamics.

Develops object-centric reinforcement learning algorithmImproves performance in complex, multi-object environmentsIntegrates actor-critic and model-based approaches

Existing state abstraction methods lack a general principle for rigorously preserving behavioral structure. This work proposes a unified framework that defines behavioral semantics in reinforcement learning in a compositional manner, grounded in local one-step descriptions of system dynamics, and establishes a theory for safe transfer of behavioral structure between abstract and concrete systems. For the first time, the framework enables a compositional formalization of behavioral semantics, supporting the derivation of quantitative metrics with correctness guarantees from logical semantics. It thus lays a principled foundation for behavioral reasoning under state abstraction and provides reusable definitions and provably faithful transfer mechanisms applicable to a broad class of behavioral structures.

behavioral semanticsbehavioral structurescompositional reasoning

Reinforcement learning (RL) faces significant bottlenecks in generalization, sample efficiency, safety, and interpretability—largely due to overreliance on low-level representation learning and insufficient integration of high-level declarative domain knowledge (e.g., facts, rules, and relational constraints). This paper presents the first systematic survey of knowledge representation and reasoning (KRR)-enhanced RL. We propose a novel logic-driven KRR-RL integration paradigm grounded in logic programming, first-order logic formalisms, differentiable logical inference, and symbolic–subsymbolic hybrid modeling. Our framework enables high-order symbolic knowledge to guide policy learning and value estimation. We categorize six core technical approaches, identify critical challenges—including knowledge grounding, scalability, and credit assignment—and outline promising future directions such as verifiable logical reasoning and dynamic knowledge acquisition.

Enhance sample efficiency and safety in RL tasks.Improve system generalization in Reinforcement Learning.Integrate Knowledge Representation for better interpretability.

Categorical semantics of compositional reinforcement learning

Aug 29, 2022
GB
Georgios Bakirtzis
🏛️ Télécom Paris | Institut Polytechnique de Paris | The University of Iowa | The University of Texas at Austin

This work addresses the challenges of modularity, interpretability, and safety in reinforcement learning (RL) task specification, with particular emphasis on compositional robustness under functional decomposition. We propose the first category-theoretic framework for RL compositionality: defining the category of Markov decision processes (MDPs), and introducing categorical constructions—including fiber products, coproducts, and pushouts—to formally characterize subtask decomposition, policy composition, and unsafe state elimination. We further pioneer the use of categorical semantics to unify modeling of state-action symmetry embeddings, sequential task concatenation, and chained task completion—represented via zig-zag diagrams. The framework rigorously establishes sufficient conditions under which “divide-and-conquer” learning yields globally optimal policies. By grounding modular RL in rigorous algebraic semantics, it enables verifiable, composable RL system design.

Characterizes minimal assumptions for robust task compositionality in RL.Develops a framework for compositional reinforcement learning representations.Unifies safety requirements and symmetries using category theory.

Latest Papers

What's happening recently
View more

This study addresses the weak theoretical foundation underlying learned representations in deep reinforcement learning, which hinders meaningful comparisons with animal learning mechanisms. Building on Markov decision process (MDP) reduction theory, the work systematically analyzes the representational structures acquired by algorithms such as DQN and PPO in navigation tasks. It reveals, for the first time, that value-based methods induce representations invariant under MDP homomorphic symmetries, whereas policy gradient methods yield representations invariant under action symmetries—a property linked to prompt dependence observed in large language models. Despite comparable task performance, these algorithms exhibit markedly different representational invariances, leading to divergent transfer capabilities. These findings offer a novel perspective on neural coding and brain-inspired learning.

deep reinforcement learninginvarianceMDP homomorphism

This work addresses the limitations of existing environment abstraction methods for large-scale Markov decision processes, which typically prioritize geometric or topological fidelity while overlooking their direct impact on policy performance. The authors propose a policy-performance-oriented tree-structured abstraction mechanism that constructs controllable approximations through state clustering and intra-cluster action distribution sharing. The abstraction granularity is dynamically adjusted based on Q-value discrepancies. Furthermore, the approach explicitly decouples value function approximation error from action-sharing loss and integrates multi-timescale reinforcement learning to enable adaptive refinement and coarsening of the abstract structure. This method achieves substantial state-space compression and demonstrates superior sample efficiency and replanning speed compared to standard actor-critic baselines.

decision-makingenvironment abstractionMarkov decision processes

This work addresses the challenge of poor generalization in real-world reinforcement learning due to large state spaces by introducing a novel approach within the CARCASS framework. For the first time, it replaces traditional Prolog with fully declarative Answer Set Programming (ASP) to construct a first-order logic-based model of Markov Decision Processes, thereby enabling logical abstraction in relational reinforcement learning. This shift significantly enhances the expressiveness and flexibility of abstract modeling. Empirical evaluations in the Blocks World and Minigrid domains demonstrate that, when augmented with domain knowledge, ASP effectively supports the construction of meaningful abstractions for reinforcement learning, confirming both its feasibility and advantages over conventional methods.

AbstractionMarkov Decision ProcessesReinforcement Learning

This work addresses the challenge of unknown policy identities in unlabeled multi-policy behavioral data by proposing Behavioral INR, a self-supervised generative model based on Implicit Neural Representations (INRs)—the first to apply INRs to unsupervised policy representation learning. The method models each policy as a function mapping states to actions and achieves policy disentanglement through segment-level latent variables and FiLM modulation. It accommodates variable-length trajectories and heterogeneous sampling granularities, and introduces a policy-level out-of-distribution (OOD) evaluation metric grounded in state-action distributions. Evaluated across diverse domains—including MuJoCo, chess, F1 racing, robotics, and Seek-Avoid—Behavioral INR substantially enhances policy identifiability in continuous state-action spaces, demonstrating superior performance particularly in settings involving long trajectories, multiple policies, and strong OOD conditions.

Behavioral Out-of-Distribution ShiftImplicit Neural RepresentationsPolicy Representation Learning

Hot Scholars

PM

Pierre Monnin

Junior Fellow in AI at Wimmics (Université Côte d'Azur, Inria, CNRS, I3S)
Semantic WebKnowledge graphsGraph EmbeddingNeurosymbolic AI
FG

Fabien Gandon

INRIA
webartificial intelligenceknowledge graphs and linked datasemantic web and ontology
CF

Catherine Faron

Professor, Univ. Côte d'Azur
Semantic WebKnowledge Representation and ReasoningOntologiesArtificial Intelligence
XL

Xiangtai Li

Research Scientist, Tiktok, SG; MMLab@NTU
Generative AIComputer Vision