Relational Object-Centric Actor-Critic

📅 2023-10-26
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing object-centric model-free and monolithic model-based methods exhibit limited generalization in image-driven reinforcement learning tasks involving numerous objects and complex dynamics. Method: We propose a novel algorithm integrating an object-centric world model with the Actor-Critic framework. Our approach jointly optimizes the world model and policy by unifying object-centric representation learning, relational inductive neural networks, and model-based RL. Crucially, it embeds predictive state/reward modeling directly into the Critic network and explicitly models actions as causal interventions on object relations—formulating model learning as causal relation induction. Contribution/Results: This is the first work to incorporate causal intervention modeling and predictive Critic design within an object-centric model-based RL framework. Experiments in 3D robotic manipulation and 2D compositional structure environments demonstrate substantial improvements over state-of-the-art object-centric model-free and monolithic model-based baselines—achieving up to 37% performance gain under high object density and intricate dynamics.
📝 Abstract
The advances in unsupervised object-centric representation learning have significantly improved its application to downstream tasks. Recent works highlight that disentangled object representations can aid policy learning in image-based, object-centric reinforcement learning tasks. This paper proposes a novel object-centric reinforcement learning algorithm that integrates actor-critic and model-based approaches by incorporating an object-centric world model within the critic. The world model captures the environment's data-generating process by predicting the next state and reward given the current state-action pair, where actions are interventions in the environment. In model-based reinforcement learning, world model learning can be interpreted as a causal induction problem, where the agent must learn the causal relationships underlying the environment's dynamics. We evaluate our method in a simulated 3D robotic environment and a 2D environment with compositional structure. As baselines, we compare against object-centric, model-free actor-critic algorithms and a state-of-the-art monolithic model-based algorithm. While the baselines show comparable performance in easier tasks, our approach outperforms them in more challenging scenarios with a large number of objects or more complex dynamics.
Problem

Research questions and friction points this paper is trying to address.

Develops object-centric reinforcement learning algorithm
Integrates actor-critic and model-based approaches
Improves performance in complex, multi-object environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates actor-critic with object-centric world model
Predicts next state and reward using causal relationships
Outperforms in complex, object-rich environments
💼 Related Jobs
No related jobs found.
FRC CSC RAS | MIPT | AIRI
L
L. Ugadiarov
FRC CSC RAS, MIPT
V
Vitaliy Vorobyov
A
A. Panov
AIRI, FRC CSC RAS