🤖 AI Summary
This study addresses the lack of convergence guarantees for Soft Actor-Critic (SAC) in continuous action spaces and clarifies the theoretical distinction between mirror descent policy objectives and classical Gibbs targets. By integrating convex optimization with reinforcement learning frameworks, we employ the Legendre differential operator to analyze Q-function curvature, rigorously proving the convergence of SAC under policy mirror descent while examining strong convexity and smoothness conditions. This work provides the first rigorous convergence guarantee for this algorithm, revealing how step sizes govern target drift and tracking error. We establish an optimal iteration complexity of O(N^{-1/5}) and demonstrate that mirror descent eliminates non-zero tracking error terms, yielding theoretically superior performance compared to conventional Gibbs-based approaches.
📝 Abstract
Soft Actor-Critic (SAC) is widely used for entropy-regularised reinforcement learning with continuous action spaces, and practical implementations perform only a few actor steps towards an evolving target. In this work, we prove convergence guarantees when the target policy arises from policy mirror descent and compare it with the classical Gibbs target. We derive sufficient conditions for the strong convexity and smoothness of the actor objective, characterised by the curvature of the $Q$-function estimate through the Legendre differential operator, and establish an $\mathcal{O}\!\left(N^{-\frac{1}{5}}\right)$ best-iterate finite-time convergence rate up to actor and critic approximation errors. Moreover, the mirror-descent step size $\lambda$ directly controls the target drift and hence actor tracking error, whereas the analogous Gibbs bound contains a non-vanishing tracking term.