apply entropy regularization

Designs and analyzes optimization objectives and training procedures that add an entropy-based penalty or bonus to encourage stochasticity and prevent probability concentration; builds entropy-regularized policy-optimization and control algorithms, derives the corresponding gradient estimators, and sets or anneals regularization schedules to trade off exploration, variance, and return.

applyentropyregularization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.65
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Average-Reward Reinforcement Learning with Entropy Regularization

Jan 15, 2025
JA
Jacob Adamczyk
🏛️ University of Massachusetts Boston | The NSF Institute for Artificial Intelligence and Fundamental Interactions | San José State University | Texas Tech University

This paper addresses the issues of discounting-induced bias and insufficient policy robustness in long-horizon decision-making tasks by proposing the first systematic entropy-regularized average-reward reinforcement learning framework. Methodologically, it tightly integrates entropy regularization with the average-reward objective, yielding a scalable algorithm based on policy gradients and dual optimization—supporting neural function approximation and online policy updates. Theoretically and technically, it fills a critical gap at the intersection of average-reward RL and entropy regularization: it eliminates the inherent temporal bias of discounted MDPs while enhancing exploration stability and resilience to environmental perturbations. Empirical evaluation on standard RL benchmarks demonstrates that the proposed algorithm consistently outperforms existing average-reward and entropy-regularized baselines across three key metrics: convergence speed, asymptotic reward performance, and policy robustness.

Entropy RegularizationLong-term PlanningReinforcement Learning

This work addresses the sensitivity of policies to environmental perturbations and the lack of theoretical robustness guarantees for entropy regularization in continuous-time reinforcement learning. For the first time, it establishes a formal connection between entropy regularization and worst-case robust reinforcement learning within continuous-time Markov decision processes, proving their equivalence to a robust optimization problem that simultaneously accounts for perturbations in both rewards and state transitions. The induced uncertainty set is explicitly characterized, shown to expand monotonically with the strength of entropy regularization, and crucially independent of action frequency. Empirical evaluations on queueing network control and market-making tasks demonstrate that entropy-regularized policies significantly outperform greedy and ε-greedy baselines under dynamic perturbations.

continuous-time reinforcement learningentropy regularizationMarkov decision processes

This work investigates the impact of entropy regularization on the convergence of policy gradient methods in stochastic exit-time control. We propose a continuous-time policy mirror descent dynamics, where entropy regularization strength is gradually annealed to enable a smooth transition from regularized solutions to the unregularized optimal policy. We establish, for the first time, a convergence rate theory for entropy-annealed mirror descent in the infinite-dimensional space of Markov kernels, revealing how entropy regularization fundamentally accelerates true gradient optimization. Theoretically, we prove exponential convergence under fixed entropy; under polynomial entropy decay, the method achieves $O(1/S)$ convergence in discrete action spaces and $O(1/sqrt{S})$ in general (continuous) action spaces. Our key innovation lies in integrating entropy annealing with infinite-dimensional variational optimization, yielding the first mirror descent framework for non-convex stochastic control problems with explicit, provable convergence rates.

Analyze continuous-time policy mirror descent with entropy annealingProve convergence rates for regularized and unregularized control problemsQuantify entropy regularization's impact on policy gradient convergence

To address the prevalent reward collapse problem in diffusion model fine-tuning, this paper proposes an entropy-regularized stochastic control framework and— for the first time—rigorously extends it to general *f*-divergence regularization. Methodologically, we formulate a continuous-time stochastic control model, integrating Itô calculus with variational inference to derive a computationally tractable and provably convergent optimal control policy. Theoretically, we establish that the proposed regularization effectively mitigates reward collapse; empirically, it significantly improves both sample quality and diversity. Key contributions include: (1) the first rigorous stochastic control analysis framework specifically designed for diffusion model fine-tuning; (2) a unified generalization of entropy regularization to arbitrary *f*-divergences, substantially enhancing methodological generality and robustness; and (3) a practical fine-tuning paradigm implementable under multiple divergence metrics.

Developing rigorous entropy-regularized fine-tuning for diffusion modelsExtending analysis to general f-divergence regularizers for fine-tuningUsing stochastic control to prevent reward collapse during generation

Fast Policy Learning for Linear Quadratic Control with Entropy Regularization

Nov 23, 2023
XG
Xin Guo
🏛️ University of California, Berkeley | University of Southern California

This work studies optimal policy learning for the infinite-horizon discounted linear quadratic control (LQC) problem with entropy regularization. To address this problem, we propose two novel algorithms: Regularized Policy Gradient (RPG) and Iterative Policy Optimization (IPO). Both algorithms achieve global linear convergence under exact policy evaluation. Moreover, IPO attains superlinear convergence within a local neighborhood and in transfer scenarios—from known to unknown environments—establishing the first superlinear convergence guarantee in LQC. To our knowledge, this is the first work to incorporate entropy regularization into the LQC policy learning framework, unifying policy gradient methods, iterative optimization, and classical linear control theory. Theoretical analysis and numerical experiments jointly validate the algorithms’ efficiency, robustness, and transferability.

Developing fast policy learning methods for entropy-regularized linear quadratic controlEnhancing policy transfer between known and unknown environment settingsProving linear and super-linear convergence rates for optimal policy discovery

Latest Papers

What's happening recently
View more

This work addresses the problem of policy synthesis in Markov decision processes (MDPs) under entropy-based constraints that enforce concentration of state visitation distributions. It formalizes entropy maximization as a policy synthesis objective for the first time, establishes its computational complexity, and introduces a novel method combining convex duality theory with invariant synthesis to handle nonlinear entropy constraints in a conditionally complete manner. By systematically analyzing the roles of memory and randomization in policies, the approach effectively synthesizes and verifies entropy-constrained policies across multiple benchmark instances, substantially extending the expressiveness and applicability of existing policy synthesis frameworks.

concentration propertycontrol policy synthesisentropy objectives

This work addresses the challenge of efficient exploration in reinforcement learning under general convex constraints—such as safety, resource limitations, or imitation requirements. It proposes the Policy Gradient Penalty (PGP) method, which incorporates state-action occupancy measure constraints via quadratic penalty functions transformed into pseudo-rewards. Under non-convex policy parameterizations, PGP provides the first global last-iterate convergence guarantee for constrained maximum entropy exploration, yielding a single policy that is provably near-optimal and nearly feasible. The theoretical analysis uncovers hidden convexity and strong duality structures inherent in the problem. Empirical evaluations in grid-world and high-dimensional continuous control tasks demonstrate the method’s effectiveness and scalability, achieving ε-optimal constrained entropy while maintaining bounded constraint violations.

constrained explorationglobal optimalitymaximum-entropy reinforcement learning

This study addresses the challenge of achieving high-probability safety in reinforcement learning under stochastic reach-avoid safety constraints. The work proposes an online algorithm that integrates entropy regularization into an optimistic framework for uncertainty (OFU) tailored to constrained Markov decision processes. Notably, this is the first approach to incorporate entropy regularization within the OFU paradigm, enabling strict safety guarantees throughout the learning process while substantially reducing inter-episode policy variability. Through finite-sample analysis, the authors derive a regret bound for the proposed algorithm and theoretically demonstrate that entropy regularization not only enhances performance but also effectively controls policy variance.

Markov decision processesonline learningsafe reinforcement learning

This work investigates the impact of the Critic on policy update variance and convergence in entropy-regularized Actor-Critic algorithms. Under a finite-horizon, discounted setting with entropy regularization, we provide the first rigorous proof that an exact Critic, when used as a baseline, substantially reduces the variance of policy gradients; moreover, even with small approximation errors, it still ensures rapid convergence. Stochastic gradient analysis reveals that, given an exact Critic, the algorithm achieves an ε-optimal regularized value function with only Õ(log(1/ε)) samples, matching the sample complexity of deterministic policy gradient methods. These findings underscore the critical importance of prioritizing accurate Critic learning in such frameworks.

actor-criticcritic estimationentropy regularization

This study addresses the continuous-time optimal asset allocation problem under stochastic volatility and portfolio constraints by introducing an entropy-regularized reinforcement learning framework. Modeling the control policy as a probability distribution rather than a deterministic function, the authors derive the associated entropy-regularized Hamilton-Jacobi-Bellman (HJB) equation via dynamic programming and propose an optimal exploratory policy of truncated Gaussian form. Leveraging stochastic control theory and martingale methods, they establish the existence of solutions to the resulting nonlinear quasilinear parabolic PDE and obtain semi-closed-form expressions for both the value function and the optimal policy. Furthermore, they develop an implementable continuous-time Actor-Critic algorithm and prove the convergence of its policy improvement process, thereby revealing an intrinsic connection between entropy-regularized relaxed controls and continuous-time reinforcement learning.

entropy regularizationoptimal investmentportfolio constraints

Hot Scholars

DT

Daniil Tiapkin

École Polytechnique
optimizationreinforcement learning
SS

Sergey Samsonov

HSE university, Moscow
high-dimensional probabilityMarkov ChainsMCMC
NM

Nikita Morozov

HSE University
Deep LearningProbabilistic ModelingGenerative ModelingReinforcement Learning