๐ค AI Summary
This study addresses the lack of finite-sample theory for distributionally robust average-reward reinforcement learning under weak communication conditions. Leveraging a distributionally robust framework and Bellman analysis techniques, we investigate the sample complexity of algorithms without prior knowledge. Core contributions include establishing explicit radius conditions under KullbackโLeibler and $f_k$-divergence balls, proving that total variation and Wasserstein balls require no weak communication assumptions, and deriving tight upper bounds on the span of the bias function. Ultimately, this work achieves polynomial sample complexity, providing rigorous finite-sample guarantees for estimating the robust optimal average reward and learning near-optimal policies. Numerical experiments further validate the $n^{-1/2}$ convergence rate.
๐ Abstract
We study distributionally robust reinforcement learning (DR-RL) in the average-reward setting under weak communication. Our main result provides finite-sample guarantees for estimating the robust optimal average reward and learning a near-optimal policy, covering both SA-rectangular and S-rectangular structures with divergence-based and distance-based uncertainty sets. Specifically, for Kullback--Leibler and $f_k$-divergence balls, we establish explicit radius conditions under which the robust average-reward Bellman equation admits a constant-gain solution, while for total variation and Wasserstein balls, any positive radius suffices without requiring the nominal MDP to be weakly communicating. Our algorithm is prior-knowledge-free and achieves sample complexities of $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-1}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$ for estimating the robust optimal average reward and $\widetilde O(|\mathcal{S}||\mathcal{A}|p_{\wedge}^{-2}\operatorname{Span}^{2}(u_{\delta}^{\ast})\epsilon^{-2})$ for learning an $\epsilon$-optimal policy. Here, $p_{\wedge}$ is the smallest positive nominal transition probability and $u_{\delta}^{\ast}$ is a robust optimal bias function. We further provide an almost-tight explicit upper bound on $\operatorname{Span}(u_{\delta}^{\ast})$. Finally, we validate the predicted $n^{-1/2}$ convergence rate through numerical experiments.