Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization

📅 2025-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses inconsistent implementations of KL regularization in Reinforcement Learning from Human Feedback (RLHF), where existing methods (e.g., GRPO) conflate the distinct functional roles of the KL term—as a reward correction versus an explicit loss. Through gradient analysis and equivalence derivation, we first prove that, under on-policy settings, optimizing “KL as a loss” is strictly gradient-equivalent to incorporating KL into the reward function. In contrast, we show that common off-policy implementations—such as adding KL as a separate loss term (e.g., $k_3$)—yield only a biased first-order approximation. To resolve this, we propose a unified framework grounded in reverse KL divergence modeling and importance sampling–based bias correction. Our theoretical analysis establishes a rigorous gradient-based foundation for KL regularization in RLHF, eliminating implementation-induced bias. Empirically, the proposed framework significantly improves training stability and sample efficiency.

Technology Category

Machine Learning: Imitation Learning & Inverse Reinforcement LearningHumans and AI: Learning Human Values and PreferencesSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Reinforcement Learning from Human Feedback (RLHF) leverages a Kullback-Leibler (KL) divergence loss to stabilize training and prevent overfitting. However, in methods such as GRPO, its implementation may be guided by principles from numerical value estimation-a practice that overlooks the term's functional role as an optimization loss. To analyze this issue, we establish a unified framework that connects two seemingly distinct implementation styles: using the mathematical term $k_n$ as a detached coefficient for the policy's score function ('$k_n$ in reward') or as a direct loss function through which gradients are propagated ('$k_n$ as loss'). We show that the latter can always be analyzed via an equivalent gradient coefficient in the former, unifying the two perspectives. Through this framework, we prove that the conventional '$k_1$ in reward' (like in PPO) is the principled loss for Reverse KL (RKL) regularization. We further establish a key finding: under on-policy conditions, the '$k_2$ as loss' formulation is, in fact, gradient-equivalent to '$k_1$ in reward'. This equivalence, first proven in our work, identifies both as the theoretically sound implementations of the RKL objective. In contrast, we show that the recently adopted '$k_3$ as loss' (like in GRPO) is merely a first-order, biased approximation of the principled loss. Furthermore, we argue that common off-policy implementations of '$k_n$ as loss' methods are biased due to neglected importance sampling, and we propose a principled correction. Our findings provide a comprehensive, gradient-based rationale for choosing and correctly implementing KL regularization, paving the way for more robust and effective RLHF systems.
Problem

Research questions and friction points this paper is trying to address.

Analyzes KL regularization implementation flaws in RLHF methods
Establishes equivalence between different KL divergence implementation styles
Proposes principled correction for biased off-policy implementations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified framework connects KL divergence implementation styles
Proves gradient equivalence between k1 reward and k2 loss
Identifies bias in GRPO method and proposes correction
K
Kezhao Liu
J
Jason Klein Liu
M
Mingtao Chen
Y
Yiming Liu