reverse kl regularization

Designs, implements, or analyzes optimization objectives and algorithms that incorporate a reverse-KL (KL divergence) regularization term or penalty—i.e., adding KL(q||p)-style or related KL regularizers into loss functions and gradient updates. Uses this construction to shape the optimization landscape and learning dynamics, for example to improve objective curvature, establish local Polyak–Łojasiewicz conditions, derive global convergence guarantees, or obtain near-linear sample-complexity bounds.

reverseklregularization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses inconsistent implementations of KL regularization in Reinforcement Learning from Human Feedback (RLHF), where existing methods (e.g., GRPO) conflate the distinct functional roles of the KL term—as a reward correction versus an explicit loss. Through gradient analysis and equivalence derivation, we first prove that, under on-policy settings, optimizing “KL as a loss” is strictly gradient-equivalent to incorporating KL into the reward function. In contrast, we show that common off-policy implementations—such as adding KL as a separate loss term (e.g., $k_3$)—yield only a biased first-order approximation. To resolve this, we propose a unified framework grounded in reverse KL divergence modeling and importance sampling–based bias correction. Our theoretical analysis establishes a rigorous gradient-based foundation for KL regularization in RLHF, eliminating implementation-induced bias. Empirically, the proposed framework significantly improves training stability and sample efficiency.

Analyzes KL regularization implementation flaws in RLHF methodsEstablishes equivalence between different KL divergence implementation stylesProposes principled correction for biased off-policy implementations

KL-Regularized Reinforcement Learning is Designed to Mode Collapse

Oct 23, 2025
AG
Anthony GX-Chen
🏛️ New York University | École Polytechnique Fédérale de Lausanne

This work addresses mode collapse in KL-regularized reinforcement learning (RL). We systematically analyze how forward and reverse KL divergences affect multimodal coverage of the target distribution, identifying regularization strength and reward scaling as key determinants of mode coverage. We propose a theoretically grounded, scalable algorithm that adaptively optimizes the target distribution solely by adjusting reward magnitude—without requiring auxiliary diversity signals. We validate our method on post-training tasks for both large language models (LLMs) and chemical language models (CLMs). Empirical results demonstrate significant improvements in generation quality and diversity under both forward- and reverse-KL settings. Crucially, our approach remains robust even under strong KL regularization or low reward scales—regimes where conventional methods fail—thereby overcoming the inherent diversity limitation of existing KL-regularized RL frameworks.

Analyzing KL divergence regularization effects in reinforcement learning optimizationDeveloping algorithm to enhance solution diversity without external signalsInvestigating mode collapse issues in language model training objectives

Statistical and Geometrical properties of regularized Kernel Kullback-Leibler divergence

Aug 29, 2024
CC
Clémentine Chazal
🏛️ CREST | ENSAE | IP Paris | INRIA | Ecole Normale Supérieure | PSL Research University

The original kernelized Kullback–Leibler (KL) divergence is ill-defined when the supports of compared distributions are disjoint—a fundamental limitation. To address this, the paper proposes a Tikhonov-regularized kernel KL divergence, constructed via covariance operator embeddings in a reproducing kernel Hilbert space (RKHS). This metric is well-defined for arbitrary probability distributions—including discrete, continuous, and mutually singular ones—and provides theoretical guarantees: a bias bound relative to the true KL divergence, finite-sample convergence rates, and a closed-form solution for discrete distributions. Furthermore, the authors formulate a Wasserstein gradient flow optimization framework for the proposed divergence, ensuring theoretical convergence, and design an efficient algorithm applicable to discrete structures such as point clouds. Experiments on point cloud transport tasks demonstrate that the method outperforms existing kernelized and Wasserstein-based approaches, achieving superior stability and robustness.

Defines regularized KKL divergence for all distributionsDerives Wasserstein gradient descent for discrete distributionsProvides bounds for regularized KKL divergence deviation

Geometry, Computation, and Optimality in Stochastic Optimization

Sep 23, 2019
CC
Chen Cheng
🏛️ Stanford University

This work systematically uncovers the decisive role of problem geometry—specifically, the curvature of the constraint set and the structure of gradients—in governing the statistical-computational trade-offs of stochastic and online optimization algorithms. We introduce the first geometric measure quantifying the deviation of a constraint set from quadratic convexity, rigorously identifying the geometric origins of suboptimality in subgradient methods. We prove that diagonal-preconditioned SGD achieves minimax-optimal convergence rates under quadratic convex constraints. For non-Euclidean, non-quadratically-convex domains—such as ℓₚ-balls with p < 2—we establish tight convergence bounds for mirror descent and adaptive gradient methods, and uncover, for the first time, a precise correspondence between their convergence rates and the accuracy-computation trade-off in Gaussian sequence estimation. Our results provide geometric criteria for algorithm selection and unify the understanding of when nonlinear updates—e.g., via mirror descent—are necessary to attain statistical optimality.

Characterize optimality of stochastic gradient methods via geometryDetermine when nonlinear updates are necessary for optimal convergenceQuantify sub-optimality of subgradient methods using constraint convexity

Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Nov 07, 2024
HZ
Heyang Zhao
🏛️ University of California, Los Angeles | University of Illinois Urbana-Champaign

This paper investigates the distinct theoretical roles of KL regularization in contextual bandits versus online Reinforcement Learning from Human Feedback (RLHF), and its interplay with data coverage. Method: We establish the first tight theoretical analysis framework for KL regularization in RLHF, introducing a two-stage hybrid sampling strategy that explicitly leverages reference policy coverage. Contribution/Results: Our analysis demonstrates that KL regularization reduces sample complexity from the standard $O(1/varepsilon^2)$ to $O(1/varepsilon)$—without requiring explicit exploration policies or strong structural assumptions on the reward function. Crucially, we identify that the data coverage induced by the reference policy directly governs online RLHF efficiency; our hybrid sampling strategy achieves sample complexity that depends additively on the coverage coefficient. The work unifies KL regularization and data coverage into a coherent theoretical framework, yielding a novel paradigm for efficient and robust human-feedback-driven policy optimization.

Analyzes KL-regularized contextual bandits and RLHFExplores data coverage impact on RLHFImproves sample complexity understanding in RLHF

Latest Papers

What's happening recently
View more

Existing methods for offline contextual bandits with forward KL regularization achieve only a slow Õ(ε⁻²) sample complexity, lacking guarantees of fast convergence. This work establishes, for the first time, a fast Õ(ε⁻¹) upper bound under the single-policy concentrability assumption by leveraging convex analysis and the principle of pessimism. We introduce a novel proof paradigm that avoids reliance on the median-of-means technique and attain this optimal rate in both tabular and general function approximation settings. Furthermore, we demonstrate that this rate is statistically tight and uncover a phenomenon wherein the slower Õ(ε⁻²) rate reemerges in the low-regularization regime.

fast ratesforward-KL regularizationoffline contextual bandits

This work addresses the ill-posedness of optimal control problems arising from traditional KL regularization, which diverges under support mismatch or in the low-noise limit. To resolve this, the authors propose a novel KL-type divergence grounded in an information-geometric framework, replacing the Fisher–Rao metric with Wasserstein and Kalman–Wasserstein transport metrics. The resulting divergence remains finite in the low-noise regime, thereby eliminating singularities and offering a geometric justification for heuristic regularization terms commonly used in Kalman ensemble methods. Through closed-form derivations for Gaussian distributions and analysis of linear time-invariant systems, the approach is validated on double-integrator and inverted pendulum systems, demonstrating both well-posedness and superior control performance.

ill-posed controlKL regularizationKullback-Leibler divergence

This work addresses the lack of a unified modular framework for analyzing adaptive optimizers, which hinders a precise characterization of their behavior under constraints on directional reachability, information budgets, and update rules. We propose a geometric–non-geometric decoupled calculus for optimizers: the geometric module, constituted by a family of positive-definite cometrics, captures realizable descent directions, while the non-geometric module governs mechanisms such as information processing, memory, and control. Within this framework, we establish a direction expressivity theorem and a residual theory for constrained cometric families, disentangling directional expressiveness from condition-number complexity and recasting optimizer design as a Pareto optimization problem under modular budgets. Theoretically, we prove that fully positive-definite geometry exactly spans all strictly descending directions; experiments demonstrate that high-information full-metric probes attain numerical precision on deterministic quadratic problems, and a Muon-style implementation preliminarily validates the auditability of matrix-operator updates.

adaptive optimizersdirection expressivitygeometric calculus

This study addresses the vulnerability of KL regularization with respect to reference policies in group-based policy optimization, systematically analyzing seven failure modes arising from its interaction with reward signals. To mitigate these issues, this work proposes Zero-Sum Calibrated Policy Optimization (ZCPO), a novel algorithm that introduces a mechanism for calibrating intra-group reward coefficients by measuring relative drift via conditional KL divergence. This calibration is further integrated into the base agent through group-relative updates, effectively circumventing the detrimental interference of KL regularization in specific scenarios. Mathematical reasoning experiments and ablation studies demonstrate that ZCPO significantly enhances both the stability and performance of policy optimization.

Failure ModesGroup Policy OptimizationKL Regularization

The statistical efficiency of regret bounds for KL-regularized multi-armed bandits under varying regularization strengths remains unclear. This work presents a refined analysis of the KL-UCB algorithm via a hierarchical peeling argument, establishing for the first time a high-probability upper bound of Õ(ηK log²T) and a non-trivial lower bound of Ω(ηK logT). These results reveal distinct scaling regimes: in the high-regularization regime, the regret scales linearly with both the regularization strength η and the number of arms K, whereas in the low-regularization regime, it transitions to an η-independent rate of Θ̃(√(KT)). The near-tight characterization fully delineates the regret behavior of KL-regularized multi-armed bandits across the regularization spectrum.

KL-regularizedmulti-armed banditsonline learning

Hot Scholars

ZC

Zehao Chen

PhD, Yale University
Porous MediaFluid DynamicsPolymerHydrogel
ZD

Zhongxiang Dai

Assistant Professor, The Chinese University of Hong Kong, Shenzhen
Machine LearningData-Centric AILarge Language ModelsMulti-Armed Bandits
XL

Xiu Li

Bytedance Seed
Computer VisionComputer Graphics3D Vision
YZ

Yifan Zhang

Princeton University
Machine LearningDeep LearningLanguage Models