🤖 AI Summary
This study addresses the vulnerability of KL regularization with respect to reference policies in group-based policy optimization, systematically analyzing seven failure modes arising from its interaction with reward signals. To mitigate these issues, this work proposes Zero-Sum Calibrated Policy Optimization (ZCPO), a novel algorithm that introduces a mechanism for calibrating intra-group reward coefficients by measuring relative drift via conditional KL divergence. This calibration is further integrated into the base agent through group-relative updates, effectively circumventing the detrimental interference of KL regularization in specific scenarios. Mathematical reasoning experiments and ablation studies demonstrate that ZCPO significantly enhances both the stability and performance of policy optimization.
📝 Abstract
Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into the base surrogate. Mathematical reasoning experiments and ablations support this design's effectiveness in our settings.