Memory-State Critic for Asymmetric Actor-Critic with Application to Vision-Based Pursuit-Evasion

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bias in policy gradient estimation arising from critics that rely solely on state or history information in partially observable Markov decision processes. To overcome this limitation, this work proposes a memory-state critic method that leverages the internal memory of the policy alongside privileged states to construct unbiased gradient estimates. Theoretically, it is proven that memory-based critics guarantee unbiasedness without requiring backpropagation through the memory module, thereby eliminating the need for recurrent approximators. Experimentally, an asymmetric actor-critic architecture is employed to validate the proposed approach in a vision-based pursuit-evasion task involving two quadrotor UAVs. Compared with history-state critics, the proposed method achieves faster convergence and superior performance while strictly preserving gradient unbiasedness.
📝 Abstract
In partially observable Markov decision processes, the optimal policy generally depends on the history of observations and past actions. Asymmetric actor-critic methods have become popular to learn such policies when additional information, such as the true state of the environment, is available during training. The critic, which is not needed at execution, is given access to the state. A critic conditioned on the state alone is generally ill-defined and yields biased policy gradients. Conditioning on the state and the history, the history-state critic restores both. In this paper, we show that conditioning the critic on the state and the policy's own memory, i.e., the internal representation of the history through which the policy selects its actions, is already well-defined and gives unbiased policy gradients, removing the need for a second recurrent approximator of the history. We call it the memory state critic. It follows that a critic based on the policy's memory need not backpropagate its loss into that memory, even though the memory is a lossy encoding of the history. We evaluate the memory-state critic in a vision-based pursuit-evasion environment between two quadrotors across two arena types. The pursuer is the learning agent, and the evader is sampled per episode from a fixed pool of heuristic behaviours. The results show that the memory-state critic outperforms the history-state critic and converges faster. In addition to being unbiased compared to the state-only critic, it maintains a slight edge in the wall arena, where the actor's history carries information that the privileged state alone does not.
Problem

Research questions and friction points this paper is trying to address.

Partially Observable Markov Decision Processes
Asymmetric Actor-Critic
Policy Gradient Bias
Memory-State Critic
Pursuit-Evasion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Memory-State Critic
Asymmetric Actor-Critic
Partially Observable Markov Decision Processes
Unbiased Policy Gradients
Vision-Based Pursuit-Evasion
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Arthur Louette
University of Liège, Belgium
A
Alejandro Sánchez Roncero
KTH Royal Institute of Technology, Sweden
G
Gaspard Lambrechts
McGill University & Mila Québec AI Institute, Canada
Pascal Leroy
Pascal Leroy
University of Liège, Belgium; Belerion, Belgium
J
Julien Hansen
University of Liège, Belgium
Petter Ögren
Petter Ögren
Professor in Computer Science and Mobile Systems, KTH (division of Robotics, Perception and Learning
RoboticsControlUnmanned Systems
Damien Ernst
Damien Ernst
Professor of Electrical Engineering and Computer Science, ULiège
Power SystemsSmart GridsReinforcement LearningEnergyMachine Learning