๐ค AI Summary
This study addresses the lack of mature finite-time and sample complexity theory for Policy Mirror Descent (PMD) in average-reward Markov Decision Processes (MDPs). By establishing a unified convergence framework based on a single master recursion, this work encompasses exact, tabular, and linear function approximation settings without requiring external regularization. The analysis reveals a superlinear convergence regime that substantially improves the dependence order on mixing time, while information-theoretic lower bounds are leveraged to demonstrate the unimprovability of the proposed algorithm. Ultimately, this work achieves an optimal sample complexity of O(t_mixยณ/ฮตยฒ), which strictly matches the theoretical lower bound. These results provide comprehensive theoretical foundations for applying PMD to average-reward MDPs.
๐ Abstract
Policy mirror descent (PMD) has a mature finite-time theory in discounted Markov decision processes (MDPs), but less is known in the average-reward setting, a more natural objective for many control applications. We give a finite-time, finite-sample analysis of PMD in ergodic average-reward MDPs built around a single master recursion that governs convergence for any critic, without external regularization. Its specializations yield linear rates for exact, inexact-tabular, and linear function approximation (LFA) updates, with a superlinear regime for exact PMD. We complement these convergence results with end-to-end sample complexities of order $t_{\mathrm{mix}}^3/\varepsilon^2$ in both tabular ($|S||A|$-dependent) and LFA ($d$-dependent) settings. Our LFA sample complexity sharpens the prior best $t_{\mathrm{mix}}^5$ mixing dependence to $t_{\mathrm{mix}}^3$, and matching information-theoretic lower bounds establish that the critic's $t_{\mathrm{mix}}^3/\varepsilon^2$ sample complexity is unimprovable in both settings.