FERPO: Forward Entropy-Regularized Policy Optimization

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of unreliable policy updates in continuous control caused by errors in the action derivatives of the critic. To overcome this limitation, it proposes a forward entropy-regularized policy optimization algorithm that directly optimizes the policy by constructing a target distribution and minimizing the forward Kullback-Leibler (KL) divergence, thereby circumventing the need for critic differentiation. By leveraging the mode-covering property of the forward KL divergence, the method effectively explores multimodal high-value regions. Furthermore, training stability is maintained through a combination of KL regularization constraints and self-normalized importance sampling. Experimental evaluations on the MuJoCo and ManiSkill benchmarks demonstrate that the proposed approach achieves competitive performance and sample efficiency, while delivering faster actor update speeds compared to REPPO.
📝 Abstract
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).
Problem

Research questions and friction points this paper is trying to address.

continuous control
online reinforcement learning
policy optimization
action gradients
critic
Innovation

Methods, ideas, or system contributions that make the work stand out.

Forward Entropy-Regularized Policy Optimization
Forward-KL objective
Self-normalized importance sampling
Maximum entropy reinforcement learning
Continuous control
🔎 Similar Papers
S
Sebastian Sanokowski
Applied and Theoretical Aspects of Robot Intelligence (ATARI) Lab, Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich
A
Alireza Sarmadi
Applied and Theoretical Aspects of Robot Intelligence (ATARI) Lab, Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich
Majid Khadiv
Majid Khadiv
Assistant Professor, TUM
Robotics