Score
Designs and analyzes policy optimization algorithms that incorporate divergence-based regularizers (e.g., f-divergences, forward/reverse KL, trust-region penalties) to constrain updates and control deviation from a reference policy. This includes constructing penalty-weighting schemes (advantage-weighted, bounded or continuous gradient weights), trust-region geometries, and update corrections to attenuate or reshape diverging policy updates.
Existing offline reinforcement learning (RL) methods predominantly employ asymmetric f-divergences—such as KL divergence—for behavioral regularization, enabling analytic policy solutions and mitigating numerical instability; symmetric f-divergences have been largely overlooked due to the absence of closed-form solutions and susceptibility to gradient explosion. Method: This paper introduces symmetric f-divergences into behavioral regularization for the first time, proposing an analytically tractable policy optimization framework based on second-order Taylor expansion. By decomposing the symmetric divergence into symmetric and conditional-symmetric components, we derive explicit closed-form policy updates and decouple the loss function to enhance numerical stability. Contribution/Results: Our approach achieves state-of-the-art performance on MuJoCo benchmarks and distribution-matching tasks, significantly outperforming mainstream offline RL algorithms. It bridges theoretical rigor—via principled symmetric regularization—with strong empirical robustness, establishing a new foundation for stable and expressive offline policy learning.
This work addresses the challenge of offline reinforcement learning under low-exploration, multi-policy datasets, where accurate value estimation and the trade-off between policy improvement and data constraints are difficult to achieve. The authors propose a novel approach that, for the first time, integrates a general and flexible $f$-divergence framework with Bellman residual constraints. By leveraging convex conjugates and linear programming, the method constructs an adaptive objective that dynamically modulates the strength of constraints imposed by complex offline data distributions. Evaluated on standard benchmarks including MuJoCo, Fetch, and AdroitHand, the approach demonstrates significant improvements in policy performance and robustly handles challenging offline datasets.
This paper investigates $f$-divergence-regularized empirical risk minimization (ERM-$f$DR) in constrained optimization settings, addressing the lack of equivalence between its solutions and explicit constraints. We first derive a dual formulation of ERM-$f$DR by leveraging the Legendre–Fenchel transform and the implicit function theorem, enabling explicit derivation of generalization error bounds without relying on strong assumptions such as strong convexity or Lipschitz continuity of loss functions. Under mild regularity conditions, we propose a generic algorithmic framework, establish an explicit generalization bound for the solution, and design an efficient method to compute the normalization constant of the regularized solution. Our core contributions are threefold: (i) unifying constrained optimization and regularization perspectives; (ii) establishing a necessary and sufficient criterion for solution–constraint equivalence; and (iii) achieving a “de-assumptionized” breakthrough in generalization analysis—removing classical structural assumptions while preserving tightness and interpretability.
To address the instability and premature convergence caused by reusing offline data in deep reinforcement learning, this paper proposes a Bregman divergence constraint mechanism grounded in state distribution. Differing from conventional approaches that define Bregman divergence over action probability spaces, our method is the first to formulate it over the space of state distributions induced by policies, thereby establishing a divergence-augmented policy optimization framework. By explicitly constraining the magnitude of policy updates’ impact on the induced state distribution, the approach ensures both safety and efficacy in offline data reuse. Evaluated on the Atari benchmark under data-scarce settings, our method significantly improves training stability and convergence speed, while achieving superior sample efficiency and policy robustness compared to mainstream algorithms including PPO and SAC. These results empirically validate the effectiveness and practicality of regularization at the state-distribution level.
This work investigates the incorporation of $f$-divergence regularization into empirical risk minimization to enhance generalization in expected risk. By establishing equivalence conditions between $f$-divergence-regularized empirical risk minimization and expected risk minimization under an $f$-divergence constraint, the study introduces the notion of a “normalizing function,” which is characterized as a nonlinear ordinary differential equation (ODE). This characterization reveals structural equivalences across different $f$-divergence regularizations. Leveraging duality theory, ODE analysis, and numerical approximation techniques, the authors develop a unified computational framework applicable to a broad class of $f$-divergences. Numerical experiments demonstrate the practical impact of various $f$-functions on training and test risks, thereby extending the range of tractable divergences and strengthening the theoretical and algorithmic coherence of the approach.
This study addresses the issue of unreliable policy updates in continuous control caused by errors in the action derivatives of the critic. To overcome this limitation, it proposes a forward entropy-regularized policy optimization algorithm that directly optimizes the policy by constructing a target distribution and minimizing the forward Kullback-Leibler (KL) divergence, thereby circumventing the need for critic differentiation. By leveraging the mode-covering property of the forward KL divergence, the method effectively explores multimodal high-value regions. Furthermore, training stability is maintained through a combination of KL regularization constraints and self-normalized importance sampling. Experimental evaluations on the MuJoCo and ManiSkill benchmarks demonstrate that the proposed approach achieves competitive performance and sample efficiency, while delivering faster actor update speeds compared to REPPO.
This work addresses the susceptibility of large language models to distributional shift in off-policy reinforcement learning, where existing approaches relying on importance sampling or hard clipping struggle to maintain stable updates under long-tailed vocabularies. The paper proposes DRPO, a novel method that refines the conventional hard-clipped trust-region mechanism by introducing a smooth, advantage-weighted quadratic regularizer based on KL divergence. This formulation preserves the geometric structure of DPPO while imposing continuous and bounded gradient weights on policy deviations, effectively mitigating divergence and providing corrective signals beyond policy boundaries. Empirical results demonstrate that DRPO significantly enhances training stability and sample efficiency across diverse model scales, architectures, and precision settings.
This study addresses the vulnerability of KL regularization with respect to reference policies in group-based policy optimization, systematically analyzing seven failure modes arising from its interaction with reward signals. To mitigate these issues, this work proposes Zero-Sum Calibrated Policy Optimization (ZCPO), a novel algorithm that introduces a mechanism for calibrating intra-group reward coefficients by measuring relative drift via conditional KL divergence. This calibration is further integrated into the base agent through group-relative updates, effectively circumventing the detrimental interference of KL regularization in specific scenarios. Mathematical reasoning experiments and ablation studies demonstrate that ZCPO significantly enhances both the stability and performance of policy optimization.
This work investigates the impact of the Critic on policy update variance and convergence in entropy-regularized Actor-Critic algorithms. Under a finite-horizon, discounted setting with entropy regularization, we provide the first rigorous proof that an exact Critic, when used as a baseline, substantially reduces the variance of policy gradients; moreover, even with small approximation errors, it still ensures rapid convergence. Stochastic gradient analysis reveals that, given an exact Critic, the algorithm achieves an ε-optimal regularized value function with only Õ(log(1/ε)) samples, matching the sample complexity of deterministic policy gradient methods. These findings underscore the critical importance of prioritizing accurate Critic learning in such frameworks.
本文通过统一正则化方法并导出性能差距的上界,提出一种约束优化问题来提高深度强化学习策略在对抗性输入扰动下的鲁棒性。