🤖 AI Summary
This study addresses the lack of theoretical guarantees that the final-iteration policy satisfies constraints in safety-critical reinforcement learning. To this end, it proposes an efficient framework for last-iterate convergence in structured constrained Markov decision processes (CMDPs), which decouples dynamics contraction from statistical error to facilitate online learning. By integrating regularized primal-dual dynamics, optimistic policy evaluation, and compactly parameterized actors, the approach accommodates both linear and general function approximation. The primary contribution lies in providing the first scalable last-iterate guarantees for large state spaces, thereby eliminating target accuracy dependence and obviating the need to maintain historical mixture policies. Furthermore, the authors derive representation-dependent complexity bounds, while experiments validate the stable convergence of the regularized method and demonstrate its superiority over unregularized baselines.
📝 Abstract
In safety-critical applications, deployment uses a single policy, whose performance and constraint satisfaction should hold directly rather than only for an average or mixture of training policies. This motivates last-iterate guarantees in constrained reinforcement learning. Recent progress has established such guarantees in exact-gradient or tabular online settings, yet scalable results for structured large-state problems remain open. We develop a general, statistically efficient framework for last-iterate convergence in structured Constrained MDPs (CMDPs). Our analysis separates contraction of the regularised primal-dual dynamics from actor approximation and statistical errors in policy evaluation under online exploration. This enables model-free on- and off-policy learning with structured function approximation: optimistic policy evaluation avoids explicit transition-model construction, while a compact parametric actor avoids maintaining mixtures or histories of past policies. We instantiate the framework for linear CMDPs and general function approximation, obtaining representation-dependent complexity and improved target-accuracy dependence over prior optimistic regularised primal-dual analyses. We further validate the stabilising effect predicted by our theory on a synthetic linear CMDP: the regularised method exhibits stable last-iterate behaviour, whereas its unregularised counterpart shows larger oscillations.