divergence regularized policy optimization

Designs and analyzes policy optimization algorithms that incorporate divergence-based regularizers (e.g., f-divergences, forward/reverse KL, trust-region penalties) to constrain updates and control deviation from a reference policy. This includes constructing penalty-weighting schemes (advantage-weighted, bounded or continuous gradient weights), trust-region geometries, and update corrections to attenuate or reshape diverging policy updates.

divergenceregularizedpolicyoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Symmetric Behavior Regularization via Taylor Expansion of Symmetry

Aug 06, 2025
LZ
Lingwei Zhu
🏛️ Great Bay University | Osaka University | University of Alberta | University of Tokyo

Existing offline reinforcement learning (RL) methods predominantly employ asymmetric f-divergences—such as KL divergence—for behavioral regularization, enabling analytic policy solutions and mitigating numerical instability; symmetric f-divergences have been largely overlooked due to the absence of closed-form solutions and susceptibility to gradient explosion. Method: This paper introduces symmetric f-divergences into behavioral regularization for the first time, proposing an analytically tractable policy optimization framework based on second-order Taylor expansion. By decomposing the symmetric divergence into symmetric and conditional-symmetric components, we derive explicit closed-form policy updates and decouple the loss function to enhance numerical stability. Contribution/Results: Our approach achieves state-of-the-art performance on MuJoCo benchmarks and distribution-matching tasks, significantly outperforming mainstream offline RL algorithms. It bridges theoretical rigor—via principled symmetric regularization—with strong empirical robustness, establishing a new foundation for stable and expressive offline policy learning.

Addresses numerical issues with symmetric divergences via Taylor expansionIntroduces symmetric divergences to offline RL frameworkProposes Symmetric f Actor-Critic for practical BRPO algorithm

This work addresses the challenge of offline reinforcement learning under low-exploration, multi-policy datasets, where accurate value estimation and the trade-off between policy improvement and data constraints are difficult to achieve. The authors propose a novel approach that, for the first time, integrates a general and flexible $f$-divergence framework with Bellman residual constraints. By leveraging convex conjugates and linear programming, the method constructs an adaptive objective that dynamically modulates the strength of constraints imposed by complex offline data distributions. Evaluated on standard benchmarks including MuJoCo, Fetch, and AdroitHand, the approach demonstrates significant improvements in policy performance and robustly handles challenging offline datasets.

Behavior Policy ConstraintDiverse Behavior Policiesf-divergence

Generalization Error of $f$-Divergence Stabilized Algorithms via Duality

Feb 20, 2025
FD
Francisco Daunas
🏛️ University of Sheffield | INRIA | Princeton University | Université de la Polynésie Française | Alan Turing Institute

This paper investigates $f$-divergence-regularized empirical risk minimization (ERM-$f$DR) in constrained optimization settings, addressing the lack of equivalence between its solutions and explicit constraints. We first derive a dual formulation of ERM-$f$DR by leveraging the Legendre–Fenchel transform and the implicit function theorem, enabling explicit derivation of generalization error bounds without relying on strong assumptions such as strong convexity or Lipschitz continuity of loss functions. Under mild regularity conditions, we propose a generic algorithmic framework, establish an explicit generalization bound for the solution, and design an efficient method to compute the normalization constant of the regularized solution. Our core contributions are threefold: (i) unifying constrained optimization and regularization perspectives; (ii) establishing a necessary and sufficient criterion for solution–constraint equivalence; and (iii) achieving a “de-assumptionized” breakthrough in generalization analysis—removing classical structural assumptions while preserving tightness and interpretability.

Characterizes generalization error for ERM-$f$DR solutionsExtends ERM-$f$DR to constrained optimization problemsIntroduces dual formulation for computational efficiency

Divergence-Augmented Policy Optimization

Jan 25, 2025
QW
Qing Wang
🏛️ Huya AI | Tencent AI Lab | The Chinese University of Hong Kong | The Hong Kong University of Science and Technology

To address the instability and premature convergence caused by reusing offline data in deep reinforcement learning, this paper proposes a Bregman divergence constraint mechanism grounded in state distribution. Differing from conventional approaches that define Bregman divergence over action probability spaces, our method is the first to formulate it over the space of state distributions induced by policies, thereby establishing a divergence-augmented policy optimization framework. By explicitly constraining the magnitude of policy updates’ impact on the induced state distribution, the approach ensures both safety and efficacy in offline data reuse. Evaluated on the Atari benchmark under data-scarce settings, our method significantly improves training stability and convergence speed, while achieving superior sample efficiency and policy robustness compared to mainstream algorithms including PPO and SAC. These results empirically validate the effectiveness and practicality of regularization at the state-distribution level.

Avoidance of Premature TerminationDeep Reinforcement LearningStable Policy Improvement

This work investigates the incorporation of $f$-divergence regularization into empirical risk minimization to enhance generalization in expected risk. By establishing equivalence conditions between $f$-divergence-regularized empirical risk minimization and expected risk minimization under an $f$-divergence constraint, the study introduces the notion of a “normalizing function,” which is characterized as a nonlinear ordinary differential equation (ODE). This characterization reveals structural equivalences across different $f$-divergence regularizations. Leveraging duality theory, ODE analysis, and numerical approximation techniques, the authors develop a unified computational framework applicable to a broad class of $f$-divergences. Numerical experiments demonstrate the practical impact of various $f$-functions on training and test risks, thereby extending the range of tractable divergences and strengthening the theoretical and algorithmic coherence of the approach.

Constraint OptimizationEmpirical Risk MinimizationExpected Risk

Latest Papers

What's happening recently
View more

This study addresses the issue of unreliable policy updates in continuous control caused by errors in the action derivatives of the critic. To overcome this limitation, it proposes a forward entropy-regularized policy optimization algorithm that directly optimizes the policy by constructing a target distribution and minimizing the forward Kullback-Leibler (KL) divergence, thereby circumventing the need for critic differentiation. By leveraging the mode-covering property of the forward KL divergence, the method effectively explores multimodal high-value regions. Furthermore, training stability is maintained through a combination of KL regularization constraints and self-normalized importance sampling. Experimental evaluations on the MuJoCo and ManiSkill benchmarks demonstrate that the proposed approach achieves competitive performance and sample efficiency, while delivering faster actor update speeds compared to REPPO.

action gradientscontinuous controlcritic

This work addresses the susceptibility of large language models to distributional shift in off-policy reinforcement learning, where existing approaches relying on importance sampling or hard clipping struggle to maintain stable updates under long-tailed vocabularies. The paper proposes DRPO, a novel method that refines the conventional hard-clipped trust-region mechanism by introducing a smooth, advantage-weighted quadratic regularizer based on KL divergence. This formulation preserves the geometric structure of DPPO while imposing continuous and bounded gradient weights on policy deviations, effectively mitigating divergence and providing corrective signals beyond policy boundaries. Empirical results demonstrate that DRPO significantly enhances training stability and sample efficiency across diverse model scales, architectures, and precision settings.

distributional shiftdivergence regularizationlarge language models

This study addresses the vulnerability of KL regularization with respect to reference policies in group-based policy optimization, systematically analyzing seven failure modes arising from its interaction with reward signals. To mitigate these issues, this work proposes Zero-Sum Calibrated Policy Optimization (ZCPO), a novel algorithm that introduces a mechanism for calibrating intra-group reward coefficients by measuring relative drift via conditional KL divergence. This calibration is further integrated into the base agent through group-relative updates, effectively circumventing the detrimental interference of KL regularization in specific scenarios. Mathematical reasoning experiments and ablation studies demonstrate that ZCPO significantly enhances both the stability and performance of policy optimization.

Failure ModesGroup Policy OptimizationKL Regularization

This work investigates the impact of the Critic on policy update variance and convergence in entropy-regularized Actor-Critic algorithms. Under a finite-horizon, discounted setting with entropy regularization, we provide the first rigorous proof that an exact Critic, when used as a baseline, substantially reduces the variance of policy gradients; moreover, even with small approximation errors, it still ensures rapid convergence. Stochastic gradient analysis reveals that, given an exact Critic, the algorithm achieves an ε-optimal regularized value function with only Õ(log(1/ε)) samples, matching the sample complexity of deterministic policy gradient methods. These findings underscore the critical importance of prioritizing accurate Critic learning in such frameworks.

actor-criticcritic estimationentropy regularization

Hot Scholars

GS

Guanya Shi

Assistant Professor, CMU RI | Amazon Scholar, FAR (Frontier AI & Robotics)
RoboticsRobot LearningReinforcement LearningControl
CS

Carmelo Sferrazza

UC Berkeley
RoboticsArtificial IntelligenceHumanoidsTactile Sensing
PA

Pieter Abbeel

UC Berkeley | Covariant
RoboticsMachine LearningAI
RD

Rocky Duan

Amazon FAR (Frontier AI & Robotics)
Robotics
DL

Dahua Lin

The Chinese University of Hong Kong
computer visionmachine learningprobabilistic inferencebayesian nonparametrics