KLip-PPO: A per-sample KL perspective on PPO-Clip

πŸ“… 2026-06-22
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the long-standing lack of theoretical unification between the two predominant variants of Proximal Policy Optimization (PPO)β€”namely, PPO with a clipped surrogate objective (PPO-Clip) and PPO with KL divergence penalty. By analyzing the per-sample KL divergence, the study reveals that PPO-Clip implicitly implements a sample-wise KL regularization with a stepwise coefficient. Through closed-form derivations involving the importance sampling ratio and the advantage function, the authors construct an adaptive KL penalty coefficient that exactly reproduces the gradient updates of PPO-Clip. Empirical validation on five MuJoCo continuous control tasks demonstrates nearly identical training curves between this reformulation and the original PPO-Clip, confirming their theoretical equivalence. This insight not only provides a unified interpretation of PPO but also opens new avenues for algorithm design based on adaptive KL regularization.
πŸ“ Abstract
Proximal Policy Optimization (PPO) is the standard policy-gradient algorithm for on-policy reinforcement learning. The literature presents it in two forms, a clipped surrogate that bounds the importance ratio between successive policies and a Kullback-Leibler penalty between them. These forms are treated as separate algorithms with their own gradients, their own hyperparameters, and their own reference implementations, and a sizeable body of empirical work compares them. We show that the gradient of the clipped surrogate is reproduced exactly by a Kullback-Leibler surrogate whose coefficient varies per sample, with closed-form dependence on the importance ratio and the advantage. The identity holds at every minibatch step and across the entire inner loop, and on five MuJoCo continuous-control benchmarks the two losses produce indistinguishable training curves. The reformulation exposes a structural feature of the clipped surrogate that the min notation hides. PPO-Clip's implicit per-sample penalty is a step function at the boundary of the trust region, and the shape of this coefficient is the natural design axis for generalising the algorithm. We sketch the resulting follow-up directions in the discussion.
Problem

Research questions and friction points this paper is trying to address.

Proximal Policy Optimization
KL divergence
PPO-Clip
importance ratio
trust region
Innovation

Methods, ideas, or system contributions that make the work stand out.

per-sample KL penalty
PPO-Clip
importance ratio
trust region
policy gradient
πŸ”Ž Similar Papers
R
Riccardo Colletti
University of California, Berkeley
R
Robin Holzinger
University of California, Berkeley