π€ AI Summary
This work addresses the long-standing lack of theoretical unification between the two predominant variants of Proximal Policy Optimization (PPO)βnamely, PPO with a clipped surrogate objective (PPO-Clip) and PPO with KL divergence penalty. By analyzing the per-sample KL divergence, the study reveals that PPO-Clip implicitly implements a sample-wise KL regularization with a stepwise coefficient. Through closed-form derivations involving the importance sampling ratio and the advantage function, the authors construct an adaptive KL penalty coefficient that exactly reproduces the gradient updates of PPO-Clip. Empirical validation on five MuJoCo continuous control tasks demonstrates nearly identical training curves between this reformulation and the original PPO-Clip, confirming their theoretical equivalence. This insight not only provides a unified interpretation of PPO but also opens new avenues for algorithm design based on adaptive KL regularization.
π Abstract
Proximal Policy Optimization (PPO) is the standard policy-gradient algorithm for on-policy reinforcement learning. The literature presents it in two forms, a clipped surrogate that bounds the importance ratio between successive policies and a Kullback-Leibler penalty between them. These forms are treated as separate algorithms with their own gradients, their own hyperparameters, and their own reference implementations, and a sizeable body of empirical work compares them. We show that the gradient of the clipped surrogate is reproduced exactly by a Kullback-Leibler surrogate whose coefficient varies per sample, with closed-form dependence on the importance ratio and the advantage. The identity holds at every minibatch step and across the entire inner loop, and on five MuJoCo continuous-control benchmarks the two losses produce indistinguishable training curves. The reformulation exposes a structural feature of the clipped surrogate that the min notation hides. PPO-Clip's implicit per-sample penalty is a step function at the boundary of the trust region, and the shape of this coefficient is the natural design axis for generalising the algorithm. We sketch the resulting follow-up directions in the discussion.