Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the minimax sample complexity of learning an ε-optimal robust policy under the average reward criterion with (s,a)-rectangular total variation uncertainty sets. Leveraging a generative model that samples from the nominal transition kernel, the work proposes a plug-in estimation–based reduction method that adaptively selects between nominal and robust discounted MDP formulations. It establishes, for the first time, a critical scale σH₀ that delineates high- and low-tolerance regimes in terms of error tolerance. The theoretical analysis reveals the structural decomposition of the sample complexity into a linear span term and a robustness-specific component, yielding matching upper and lower bounds. The proposed algorithm is shown to achieve near-optimal performance—within logarithmic factors—and demonstrates empirical efficacy across varying levels of error tolerance.
📝 Abstract
Distributionally robust Markov decision processes provide a principled framework for sequential decision making under model uncertainty. We study how many samples are necessary and sufficient to learn an $\varepsilon$-optimal robust policy under the average-reward criterion. A generative model provides samples from the nominal transition kernel, whereas policy performance is evaluated over $(s,a)$-rectangular total-variation uncertainty sets of radius at most $σ$. Let $H_0$ and $H_σ$ denote the nominal and robust optimal bias spans, respectively. We identify $σH_0$ as the perturbation scale separating high- and low-tolerance regimes. Our matching upper and lower bounds show that, up to logarithmic factors, the minimax total sample complexity is $$ NSA \asymp \frac{SA}{\varepsilon^2}\begin{cases} \min\{H_0,H_σ\}, & \varepsilon\gtrsimσH_0,\\ \min\{H_0,H_σ\}+σH_σ^2, & \varepsilon\lesssimσH_0. \end{cases} $$ Here $S$ and $A$ are the numbers of states and actions, and $N$ is the number of samples per state-action pair. The sample complexity consists of a linear-span term that resembles the nominal AMDP results and a robustness-specific term that appears only in the low-tolerance regime. We attain these rates using reduction-based plug-in procedures that select the reduction---nominal or robust---and its discount factor: a span-informed procedure that makes these choices using known span parameters, and a span-agnostic procedure that calibrates both choices from data.
Problem

Research questions and friction points this paper is trying to address.

Robust Markov Decision Processes
Average-Reward Criterion
Sample Complexity
Model Uncertainty
Distributional Robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

distributionally robust MDP
average-reward criterion
minimax sample complexity
plug-in reduction
bias span