primal-dual sac

Designs and implements primal-dual Soft Actor-Critic (SAC) algorithms that incorporate Lagrangian dual variables to enforce constraints by jointly optimizing the stochastic actor, entropy-regularized critic(s), and dual parameters. Builds and analyzes constrained SAC variants (e.g., multi-head cost critics, constrained SAC, primal-dual SAC) to satisfy average or expected constraints while preserving sample efficiency and enabling single-forward-pass inference for real-time operation.

primal-dualsac

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.48
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

DSAC-C: Constrained Maximum Entropy for Robust Discrete Soft-Actor Critic

Oct 26, 2023
DN
Dexter Neo
🏛️ National University of Singapore

Soft Actor-Critic (SAC) for discrete action spaces suffers from weak in-distribution and out-of-distribution generalization, low sample efficiency, and insufficient robustness for safe deployment. Method: We propose a maximum-entropy reinforcement learning framework with statistical constraints, introducing— for the first time in discrete SAC—a proxy critic-guided statistical regularization mechanism that explicitly enforces distributional robustness of the policy entropy objective, thereby mitigating domain shift effects. Contribution/Results: Evaluated on Atari 2600 under low-data regimes, our method achieves an average performance gain of 12.7% over baselines on both in-distribution and out-of-distribution tasks. It significantly improves policy generalization and deployment robustness, establishing a novel paradigm for discrete control under real-world constraints—namely, limited samples and dynamically shifting environments.

Enhancing discrete SAC with Maximum Entropy constraintsImproving robustness against domain shifts in RLValidating performance in low-data Atari game scenarios

This study addresses the lack of convergence guarantees for Soft Actor-Critic (SAC) in continuous action spaces and clarifies the theoretical distinction between mirror descent policy objectives and classical Gibbs targets. By integrating convex optimization with reinforcement learning frameworks, we employ the Legendre differential operator to analyze Q-function curvature, rigorously proving the convergence of SAC under policy mirror descent while examining strong convexity and smoothness conditions. This work provides the first rigorous convergence guarantee for this algorithm, revealing how step sizes govern target drift and tracking error. We establish an optimal iteration complexity of O(N^{-1/5}) and demonstrate that mirror descent eliminates non-zero tracking error terms, yielding theoretically superior performance compared to conventional Gibbs-based approaches.

Convergence GuaranteesEntropy RegularizationMirror Descent

Standard Soft Actor-Critic (SAC) employs reverse KL divergence for policy updates, rendering the optimal policy projection analytically intractable and necessitating gradient-based approximations—leading to training instability and poor sample efficiency. This work proposes forward KL divergence as a principled alternative, enabling the first closed-form optimal policy projection within SAC. We further introduce a bidirectional optimization framework integrating forward initialization with reverse fine-tuning. Theoretically, our approach unifies KL analysis under Gaussian policies, Boltzmann action marginal modeling, and policy projection theory; methodologically, it jointly ensures stability and optimality. Evaluated on standard continuous-control benchmarks, the proposed method achieves a 30% average improvement in episode return, while significantly enhancing sample efficiency and training robustness.

Improves sample efficiency and episodic rewards by 30%Investigates forward KL divergence in SAC for policy updatesProposes Bidirectional SAC combining forward and reverse KL advantages

SACn: Soft Actor-Critic with n-step Returns

Dec 15, 2025
JŁ
Jakub Łyskawa
🏛️ Warsaw University of Technology | Warsaw University | IDEAS Research Institute

Direct integration of Soft Actor-Critic (SAC) with n-step returns introduces off-policy bias due to policy drift, while conventional importance sampling suffers from numerical instability and high variance. This paper proposes SACn, the first method enabling safe and stable fusion of SAC with n-step entropy-regularized reinforcement learning. Its core innovations are: (1) τ-sampled entropy estimation, which reduces variance in the target Q-function by decoupling entropy estimation from policy evaluation; and (2) a simplified importance sampling mechanism that eliminates high-variance weight accumulation and alleviates hyperparameter sensitivity. Evaluated on the MuJoCo benchmark, SACn achieves significantly faster convergence and improved policy stability compared to standard SAC and multiple n-step baselines. It consistently outperforms these methods across diverse tasks, establishing a robust and efficient framework for off-policy n-step maximum-entropy RL.

Addresses bias and instability in off-policy n-step importance samplingCombines SAC with n-step returns to increase convergence speedReduces variance in entropy estimation for stable learning

Existing model-free reinforcement learning lacks a general-purpose mechanism for enforcing policy constraints; prior approaches support only specific constraint types (e.g., value or density constraints). Method: We propose DualCRL, a unified primal-dual framework that models diverse behavioral constraints—including action density bounds and state-action transition costs—as learnable reward shaping terms, seamlessly integrating with both value-based and actor-critic paradigms. Contribution/Results: DualCRL establishes, for the first time, the formal equivalence between dual variables and reward shaping; introduces novel constraint classes; and enables end-to-end joint optimization of multiple heterogeneous constraints. Empirical evaluation in interpretable environments demonstrates significantly improved training stability under multi-constraint settings and yields a plug-and-play policy constraint toolkit.

Imposing behavioral constraints on model-free RL policiesIntroducing novel constraints like action density boundsUnifying existing techniques with primal-dual framework

Latest Papers

What's happening recently
View more

This work addresses the challenge of simultaneously satisfying multiple objectives and safety constraints in infinite-horizon average-reward reinforcement learning. The authors propose a multi-level Monte Carlo (MLMC)-based primal-dual natural actor-critic algorithm that, for the first time, handles gradient estimation bias induced by nonlinear scalarization in the average-reward setting—without requiring prior knowledge of mixing times. The method jointly controls bias in the objective function, constraint evaluation, and policy gradient estimates. The algorithm achieves the state-of-the-art theoretical guarantees with a global convergence rate and constraint violation bound both scaling as $\tilde{O}(1/\sqrt{T})$, thereby attaining optimal sample complexity for multi-objective safe reinforcement learning with nonlinear scalarization.

average-reward RLbias in policy gradientsmulti-objective optimization

This work investigates the impact of the Critic on policy update variance and convergence in entropy-regularized Actor-Critic algorithms. Under a finite-horizon, discounted setting with entropy regularization, we provide the first rigorous proof that an exact Critic, when used as a baseline, substantially reduces the variance of policy gradients; moreover, even with small approximation errors, it still ensures rapid convergence. Stochastic gradient analysis reveals that, given an exact Critic, the algorithm achieves an ε-optimal regularized value function with only Õ(log(1/ε)) samples, matching the sample complexity of deterministic policy gradient methods. These findings underscore the critical importance of prioritizing accurate Critic learning in such frameworks.

actor-criticcritic estimationentropy regularization

This work addresses the poor scalability of Soft Actor-Critic (SAC) in large-scale parallel training, which has hindered its application to high-performance legged robot control compared to Proximal Policy Optimization (PPO). By introducing three key improvements—optimized policy initialization, timeout-aware critic targets, and multi-step return estimation—the proposed method substantially enhances SAC’s training stability and sample efficiency in massively parallel environments. For the first time, SAC achieves performance on par with PPO across diverse legged robot platforms and a wide range of locomotion tasks. Furthermore, the approach enables efficient and robust simulation-to-reality (sim-to-real) transfer, effectively overcoming a longstanding barrier that has limited off-policy algorithms in large-scale simulation and real-world online learning for legged robotics.

legged locomotionmassively parallel trainingsample efficiency

This work addresses the challenges of exploration and inefficient policy learning in sparse-reward, long-horizon tasks by proposing a two-level hierarchical reinforcement learning framework. The high-level controller performs strategic planning to guide long-term exploration, while the low-level policy leverages Soft Actor-Critic (SAC) for continuous control, augmented with entropy regularization to enhance both policy diversity and stability. By effectively integrating hierarchical structure with maximum-entropy learning, the proposed method significantly outperforms standard SAC baselines on the SAR-2 dataset, achieving notable improvements in task success rate, environmental coverage efficiency, and convergence speed.

continuous controlexplorationlong-horizon

Hot Scholars

ZL

Zhankun Luo

Purdue University
statistical learninggenerative diffusion modelimage processingdeep learning
AH

Abolfazl Hashemi

Assistant Professor of ECE, Purdue University
Large-Scale Optimization
MX

Maosheng Xiong

Associate Professor of Mathematics, Hong Kong University of Science and Technology
Number theorycoding theoryinformation theory