Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of theoretical guarantees that the final-iteration policy satisfies constraints in safety-critical reinforcement learning. To this end, it proposes an efficient framework for last-iterate convergence in structured constrained Markov decision processes (CMDPs), which decouples dynamics contraction from statistical error to facilitate online learning. By integrating regularized primal-dual dynamics, optimistic policy evaluation, and compactly parameterized actors, the approach accommodates both linear and general function approximation. The primary contribution lies in providing the first scalable last-iterate guarantees for large state spaces, thereby eliminating target accuracy dependence and obviating the need to maintain historical mixture policies. Furthermore, the authors derive representation-dependent complexity bounds, while experiments validate the stable convergence of the regularized method and demonstrate its superiority over unregularized baselines.
📝 Abstract
In safety-critical applications, deployment uses a single policy, whose performance and constraint satisfaction should hold directly rather than only for an average or mixture of training policies. This motivates last-iterate guarantees in constrained reinforcement learning. Recent progress has established such guarantees in exact-gradient or tabular online settings, yet scalable results for structured large-state problems remain open. We develop a general, statistically efficient framework for last-iterate convergence in structured Constrained MDPs (CMDPs). Our analysis separates contraction of the regularised primal-dual dynamics from actor approximation and statistical errors in policy evaluation under online exploration. This enables model-free on- and off-policy learning with structured function approximation: optimistic policy evaluation avoids explicit transition-model construction, while a compact parametric actor avoids maintaining mixtures or histories of past policies. We instantiate the framework for linear CMDPs and general function approximation, obtaining representation-dependent complexity and improved target-accuracy dependence over prior optimistic regularised primal-dual analyses. We further validate the stabilising effect predicted by our theory on a synthetic linear CMDP: the regularised method exhibits stable last-iterate behaviour, whereas its unregularised counterpart shows larger oscillations.
Problem

Research questions and friction points this paper is trying to address.

Constrained Reinforcement Learning
Last-Iterate Convergence
Constrained MDPs
Function Approximation
Safety-Critical Applications
Innovation

Methods, ideas, or system contributions that make the work stand out.

Last-Iterate Convergence
Constrained MDPs
Primal-Dual Dynamics
Function Approximation
Model-Free Reinforcement Learning
N
Nam Phuong Tran
Inria centre at the University of Lille
T
Trinh Ha Mai Huynh
Vin Motion
T
Tuyen Pham Le
Vin Motion
V
Van-Truong Nguyen
Vin Motion
Q
Quan Nguyen
Vin Motion
Long Tran-Thanh
Long Tran-Thanh
Professor in Computer Science, University of Warwick
Artificial IntelligenceAI for social goodgame theoryhuman-agent learningmulti-armed bandits