Score
Designs and implements primal-dual Soft Actor-Critic (SAC) algorithms that incorporate Lagrangian dual variables to enforce constraints by jointly optimizing the stochastic actor, entropy-regularized critic(s), and dual parameters. Builds and analyzes constrained SAC variants (e.g., multi-head cost critics, constrained SAC, primal-dual SAC) to satisfy average or expected constraints while preserving sample efficiency and enabling single-forward-pass inference for real-time operation.
Soft Actor-Critic (SAC) for discrete action spaces suffers from weak in-distribution and out-of-distribution generalization, low sample efficiency, and insufficient robustness for safe deployment. Method: We propose a maximum-entropy reinforcement learning framework with statistical constraints, introducing— for the first time in discrete SAC—a proxy critic-guided statistical regularization mechanism that explicitly enforces distributional robustness of the policy entropy objective, thereby mitigating domain shift effects. Contribution/Results: Evaluated on Atari 2600 under low-data regimes, our method achieves an average performance gain of 12.7% over baselines on both in-distribution and out-of-distribution tasks. It significantly improves policy generalization and deployment robustness, establishing a novel paradigm for discrete control under real-world constraints—namely, limited samples and dynamically shifting environments.
This study addresses the lack of convergence guarantees for Soft Actor-Critic (SAC) in continuous action spaces and clarifies the theoretical distinction between mirror descent policy objectives and classical Gibbs targets. By integrating convex optimization with reinforcement learning frameworks, we employ the Legendre differential operator to analyze Q-function curvature, rigorously proving the convergence of SAC under policy mirror descent while examining strong convexity and smoothness conditions. This work provides the first rigorous convergence guarantee for this algorithm, revealing how step sizes govern target drift and tracking error. We establish an optimal iteration complexity of O(N^{-1/5}) and demonstrate that mirror descent eliminates non-zero tracking error terms, yielding theoretically superior performance compared to conventional Gibbs-based approaches.
Standard Soft Actor-Critic (SAC) employs reverse KL divergence for policy updates, rendering the optimal policy projection analytically intractable and necessitating gradient-based approximations—leading to training instability and poor sample efficiency. This work proposes forward KL divergence as a principled alternative, enabling the first closed-form optimal policy projection within SAC. We further introduce a bidirectional optimization framework integrating forward initialization with reverse fine-tuning. Theoretically, our approach unifies KL analysis under Gaussian policies, Boltzmann action marginal modeling, and policy projection theory; methodologically, it jointly ensures stability and optimality. Evaluated on standard continuous-control benchmarks, the proposed method achieves a 30% average improvement in episode return, while significantly enhancing sample efficiency and training robustness.
Direct integration of Soft Actor-Critic (SAC) with n-step returns introduces off-policy bias due to policy drift, while conventional importance sampling suffers from numerical instability and high variance. This paper proposes SACn, the first method enabling safe and stable fusion of SAC with n-step entropy-regularized reinforcement learning. Its core innovations are: (1) τ-sampled entropy estimation, which reduces variance in the target Q-function by decoupling entropy estimation from policy evaluation; and (2) a simplified importance sampling mechanism that eliminates high-variance weight accumulation and alleviates hyperparameter sensitivity. Evaluated on the MuJoCo benchmark, SACn achieves significantly faster convergence and improved policy stability compared to standard SAC and multiple n-step baselines. It consistently outperforms these methods across diverse tasks, establishing a robust and efficient framework for off-policy n-step maximum-entropy RL.
Existing model-free reinforcement learning lacks a general-purpose mechanism for enforcing policy constraints; prior approaches support only specific constraint types (e.g., value or density constraints). Method: We propose DualCRL, a unified primal-dual framework that models diverse behavioral constraints—including action density bounds and state-action transition costs—as learnable reward shaping terms, seamlessly integrating with both value-based and actor-critic paradigms. Contribution/Results: DualCRL establishes, for the first time, the formal equivalence between dual variables and reward shaping; introduces novel constraint classes; and enables end-to-end joint optimization of multiple heterogeneous constraints. Empirical evaluation in interpretable environments demonstrates significantly improved training stability under multi-constraint settings and yields a plug-and-play policy constraint toolkit.
This work addresses the challenge of simultaneously satisfying multiple objectives and safety constraints in infinite-horizon average-reward reinforcement learning. The authors propose a multi-level Monte Carlo (MLMC)-based primal-dual natural actor-critic algorithm that, for the first time, handles gradient estimation bias induced by nonlinear scalarization in the average-reward setting—without requiring prior knowledge of mixing times. The method jointly controls bias in the objective function, constraint evaluation, and policy gradient estimates. The algorithm achieves the state-of-the-art theoretical guarantees with a global convergence rate and constraint violation bound both scaling as $\tilde{O}(1/\sqrt{T})$, thereby attaining optimal sample complexity for multi-objective safe reinforcement learning with nonlinear scalarization.
This work investigates the impact of the Critic on policy update variance and convergence in entropy-regularized Actor-Critic algorithms. Under a finite-horizon, discounted setting with entropy regularization, we provide the first rigorous proof that an exact Critic, when used as a baseline, substantially reduces the variance of policy gradients; moreover, even with small approximation errors, it still ensures rapid convergence. Stochastic gradient analysis reveals that, given an exact Critic, the algorithm achieves an ε-optimal regularized value function with only Õ(log(1/ε)) samples, matching the sample complexity of deterministic policy gradient methods. These findings underscore the critical importance of prioritizing accurate Critic learning in such frameworks.
研究通过单循环、熵正则化的自然演员-评论家算法解决了在无正则化目标下加速收敛的问题,利用指数平移机制实现了更快的收敛速度。
This work addresses the poor scalability of Soft Actor-Critic (SAC) in large-scale parallel training, which has hindered its application to high-performance legged robot control compared to Proximal Policy Optimization (PPO). By introducing three key improvements—optimized policy initialization, timeout-aware critic targets, and multi-step return estimation—the proposed method substantially enhances SAC’s training stability and sample efficiency in massively parallel environments. For the first time, SAC achieves performance on par with PPO across diverse legged robot platforms and a wide range of locomotion tasks. Furthermore, the approach enables efficient and robust simulation-to-reality (sim-to-real) transfer, effectively overcoming a longstanding barrier that has limited off-policy algorithms in large-scale simulation and real-world online learning for legged robotics.
This work addresses the challenges of exploration and inefficient policy learning in sparse-reward, long-horizon tasks by proposing a two-level hierarchical reinforcement learning framework. The high-level controller performs strategic planning to guide long-term exploration, while the low-level policy leverages Soft Actor-Critic (SAC) for continuous control, augmented with entropy regularization to enhance both policy diversity and stability. By effectively integrating hierarchical structure with maximum-entropy learning, the proposed method significantly outperforms standard SAC baselines on the SAR-2 dataset, achieving notable improvements in task success rate, environmental coverage efficiency, and convergence speed.