Fast Regularized Policy Mirror Descent with One-Step TD Updates

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the slow convergence of regularized policy mirror descent caused by its reliance on exact evaluation, and proposes a stochastic optimization algorithm incorporating one-step temporal difference (TD) updates. The proposed method eliminates the need for trajectory resets or nested evaluations by introducing resolvent auxiliary distributions, Bregman divergence bounds, and visitation-weighted estimation techniques to achieve efficient computation. Theoretical analysis establishes that the algorithm attains a global linear convergence rate along with an inverse-linear sample complexity guarantee. Numerical experiments further validate the effectiveness of these theoretical results.
📝 Abstract
Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or increasingly accurate policy evaluation. We analyze PMD coupled with a persistent critic advanced by one temporal-difference (TD) update. For finite discounted MDPs, we establish global linear convergence in value for exact coordinate-wise Bellman updates, with any positive constant actor stepsize and arbitrary finite critic initialization. The proof combines a resolvent-based auxiliary distribution with a decaying Bellman-violation correction and a potential weighted by inverse coordinate weights. We then study stochastic TD-PMD with general strongly convex mirror maps under a single off-policy Markov trajectory. With suitably chosen constant stepsizes and a finite-batch TD update, the method achieves an expected value gap of $ε$ after $\widetilde{O}(1/((1-γ)^5 \widetildeσ_b ε))$ transitions. The stochastic analysis relies on the trajectory-wise Lipschitz continuity of the regularizer, derived from uniform bounds on vertex Bregman divergences, together with a visitation-weighted resolvent estimate for signed critic-error propagation that yields an inverse-linear dependence on behavior coverage $\widetildeσ_b$. In contrast to many prior guarantees for regularized policy optimization, our sample-complexity guarantee holds without trajectory resets, generative-model access, or nested policy-evaluation loops. Numerical results are consistent with the theoretical convergence analysis.
Problem

Research questions and friction points this paper is trying to address.

Policy Mirror Descent
Temporal-Difference Learning
Regularized MDPs
Sample Complexity
Convergence Analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy Mirror Descent
One-Step TD Updates
Linear Convergence
Sample Complexity
Off-Policy Reinforcement Learning
🔎 Similar Papers