Score
Designs and analyzes optimization objectives and training procedures that add an entropy-based penalty or bonus to encourage stochasticity and prevent probability concentration; builds entropy-regularized policy-optimization and control algorithms, derives the corresponding gradient estimators, and sets or anneals regularization schedules to trade off exploration, variance, and return.
This paper addresses the issues of discounting-induced bias and insufficient policy robustness in long-horizon decision-making tasks by proposing the first systematic entropy-regularized average-reward reinforcement learning framework. Methodologically, it tightly integrates entropy regularization with the average-reward objective, yielding a scalable algorithm based on policy gradients and dual optimization—supporting neural function approximation and online policy updates. Theoretically and technically, it fills a critical gap at the intersection of average-reward RL and entropy regularization: it eliminates the inherent temporal bias of discounted MDPs while enhancing exploration stability and resilience to environmental perturbations. Empirical evaluation on standard RL benchmarks demonstrates that the proposed algorithm consistently outperforms existing average-reward and entropy-regularized baselines across three key metrics: convergence speed, asymptotic reward performance, and policy robustness.
This work addresses the sensitivity of policies to environmental perturbations and the lack of theoretical robustness guarantees for entropy regularization in continuous-time reinforcement learning. For the first time, it establishes a formal connection between entropy regularization and worst-case robust reinforcement learning within continuous-time Markov decision processes, proving their equivalence to a robust optimization problem that simultaneously accounts for perturbations in both rewards and state transitions. The induced uncertainty set is explicitly characterized, shown to expand monotonically with the strength of entropy regularization, and crucially independent of action frequency. Empirical evaluations on queueing network control and market-making tasks demonstrate that entropy-regularized policies significantly outperform greedy and ε-greedy baselines under dynamic perturbations.
This work investigates the impact of entropy regularization on the convergence of policy gradient methods in stochastic exit-time control. We propose a continuous-time policy mirror descent dynamics, where entropy regularization strength is gradually annealed to enable a smooth transition from regularized solutions to the unregularized optimal policy. We establish, for the first time, a convergence rate theory for entropy-annealed mirror descent in the infinite-dimensional space of Markov kernels, revealing how entropy regularization fundamentally accelerates true gradient optimization. Theoretically, we prove exponential convergence under fixed entropy; under polynomial entropy decay, the method achieves $O(1/S)$ convergence in discrete action spaces and $O(1/sqrt{S})$ in general (continuous) action spaces. Our key innovation lies in integrating entropy annealing with infinite-dimensional variational optimization, yielding the first mirror descent framework for non-convex stochastic control problems with explicit, provable convergence rates.
To address the prevalent reward collapse problem in diffusion model fine-tuning, this paper proposes an entropy-regularized stochastic control framework and— for the first time—rigorously extends it to general *f*-divergence regularization. Methodologically, we formulate a continuous-time stochastic control model, integrating Itô calculus with variational inference to derive a computationally tractable and provably convergent optimal control policy. Theoretically, we establish that the proposed regularization effectively mitigates reward collapse; empirically, it significantly improves both sample quality and diversity. Key contributions include: (1) the first rigorous stochastic control analysis framework specifically designed for diffusion model fine-tuning; (2) a unified generalization of entropy regularization to arbitrary *f*-divergences, substantially enhancing methodological generality and robustness; and (3) a practical fine-tuning paradigm implementable under multiple divergence metrics.
This work studies optimal policy learning for the infinite-horizon discounted linear quadratic control (LQC) problem with entropy regularization. To address this problem, we propose two novel algorithms: Regularized Policy Gradient (RPG) and Iterative Policy Optimization (IPO). Both algorithms achieve global linear convergence under exact policy evaluation. Moreover, IPO attains superlinear convergence within a local neighborhood and in transfer scenarios—from known to unknown environments—establishing the first superlinear convergence guarantee in LQC. To our knowledge, this is the first work to incorporate entropy regularization into the LQC policy learning framework, unifying policy gradient methods, iterative optimization, and classical linear control theory. Theoretical analysis and numerical experiments jointly validate the algorithms’ efficiency, robustness, and transferability.
This work addresses the problem of policy synthesis in Markov decision processes (MDPs) under entropy-based constraints that enforce concentration of state visitation distributions. It formalizes entropy maximization as a policy synthesis objective for the first time, establishes its computational complexity, and introduces a novel method combining convex duality theory with invariant synthesis to handle nonlinear entropy constraints in a conditionally complete manner. By systematically analyzing the roles of memory and randomization in policies, the approach effectively synthesizes and verifies entropy-constrained policies across multiple benchmark instances, substantially extending the expressiveness and applicability of existing policy synthesis frameworks.
This work addresses the challenge of efficient exploration in reinforcement learning under general convex constraints—such as safety, resource limitations, or imitation requirements. It proposes the Policy Gradient Penalty (PGP) method, which incorporates state-action occupancy measure constraints via quadratic penalty functions transformed into pseudo-rewards. Under non-convex policy parameterizations, PGP provides the first global last-iterate convergence guarantee for constrained maximum entropy exploration, yielding a single policy that is provably near-optimal and nearly feasible. The theoretical analysis uncovers hidden convexity and strong duality structures inherent in the problem. Empirical evaluations in grid-world and high-dimensional continuous control tasks demonstrate the method’s effectiveness and scalability, achieving ε-optimal constrained entropy while maintaining bounded constraint violations.
This study addresses the challenge of achieving high-probability safety in reinforcement learning under stochastic reach-avoid safety constraints. The work proposes an online algorithm that integrates entropy regularization into an optimistic framework for uncertainty (OFU) tailored to constrained Markov decision processes. Notably, this is the first approach to incorporate entropy regularization within the OFU paradigm, enabling strict safety guarantees throughout the learning process while substantially reducing inter-episode policy variability. Through finite-sample analysis, the authors derive a regret bound for the proposed algorithm and theoretically demonstrate that entropy regularization not only enhances performance but also effectively controls policy variance.
This work investigates the impact of the Critic on policy update variance and convergence in entropy-regularized Actor-Critic algorithms. Under a finite-horizon, discounted setting with entropy regularization, we provide the first rigorous proof that an exact Critic, when used as a baseline, substantially reduces the variance of policy gradients; moreover, even with small approximation errors, it still ensures rapid convergence. Stochastic gradient analysis reveals that, given an exact Critic, the algorithm achieves an ε-optimal regularized value function with only Õ(log(1/ε)) samples, matching the sample complexity of deterministic policy gradient methods. These findings underscore the critical importance of prioritizing accurate Critic learning in such frameworks.
This study addresses the continuous-time optimal asset allocation problem under stochastic volatility and portfolio constraints by introducing an entropy-regularized reinforcement learning framework. Modeling the control policy as a probability distribution rather than a deterministic function, the authors derive the associated entropy-regularized Hamilton-Jacobi-Bellman (HJB) equation via dynamic programming and propose an optimal exploratory policy of truncated Gaussian form. Leveraging stochastic control theory and martingale methods, they establish the existence of solutions to the resulting nonlinear quasilinear parabolic PDE and obtain semi-closed-form expressions for both the value function and the optimal policy. Furthermore, they develop an implementable continuous-time Actor-Critic algorithm and prove the convergence of its policy improvement process, thereby revealing an intrinsic connection between entropy-regularized relaxed controls and continuous-time reinforcement learning.