Sharp Convergence and Sample Complexity of Policy Mirror Descent for Average-Reward MDPs

๐Ÿ“… 2026-10-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the lack of mature finite-time and sample complexity theory for Policy Mirror Descent (PMD) in average-reward Markov Decision Processes (MDPs). By establishing a unified convergence framework based on a single master recursion, this work encompasses exact, tabular, and linear function approximation settings without requiring external regularization. The analysis reveals a superlinear convergence regime that substantially improves the dependence order on mixing time, while information-theoretic lower bounds are leveraged to demonstrate the unimprovability of the proposed algorithm. Ultimately, this work achieves an optimal sample complexity of O(t_mixยณ/ฮตยฒ), which strictly matches the theoretical lower bound. These results provide comprehensive theoretical foundations for applying PMD to average-reward MDPs.
๐Ÿ“ Abstract
Policy mirror descent (PMD) has a mature finite-time theory in discounted Markov decision processes (MDPs), but less is known in the average-reward setting, a more natural objective for many control applications. We give a finite-time, finite-sample analysis of PMD in ergodic average-reward MDPs built around a single master recursion that governs convergence for any critic, without external regularization. Its specializations yield linear rates for exact, inexact-tabular, and linear function approximation (LFA) updates, with a superlinear regime for exact PMD. We complement these convergence results with end-to-end sample complexities of order $t_{\mathrm{mix}}^3/\varepsilon^2$ in both tabular ($|S||A|$-dependent) and LFA ($d$-dependent) settings. Our LFA sample complexity sharpens the prior best $t_{\mathrm{mix}}^5$ mixing dependence to $t_{\mathrm{mix}}^3$, and matching information-theoretic lower bounds establish that the critic's $t_{\mathrm{mix}}^3/\varepsilon^2$ sample complexity is unimprovable in both settings.
Problem

Research questions and friction points this paper is trying to address.

Policy Mirror Descent
Average-Reward MDPs
Sample Complexity
Convergence Analysis
Ergodic MDPs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy Mirror Descent
Average-Reward MDPs
Sample Complexity
Linear Function Approximation
Convergence Rate
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.