Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the theoretical bottleneck in weakly coupled Markov decision processes (MDPs), where the optimality gap of existing policies is fundamentally constrained to $O(1/\sqrt{N})$. Focusing on average-reward weakly coupled MDPs, this work integrates mean-field game theory with local linear dynamics modeling to propose a non-prioritized policy design paradigm. By transcending the limitations of traditional priority-ranking frameworks, the proposed approach effectively induces locally linear mean-field dynamic evolution within the system. The primary contribution of this research is the first significant tightening of the optimality gap from $O(1/\sqrt{N})$ to $O(1/N)$, substantially surpassing established theoretical bounds. This advancement provides a novel theoretical foundation for the near-optimal control of large-scale multi-agent systems.
📝 Abstract
We study average-reward weakly-coupled Markov decision processes (WCMDPs), where a WCMDP consists of $N$ smaller MDPs, called arms, that share multiple per-step budget constraints. We consider the setting where the arms have identical model parameters, multiple actions, and state- and action-dependent costs. For restless bandits (RBs), a well-studied special case of WCMDPs, prior work has developed policies that achieve an $O(1/\sqrt{N})$ optimality gap under general conditions, and has further identified conditions under which policies can achieve a better-than-$1/\sqrt{N}$ optimality gap. However, for general WCMDPs, no prior result achieves an optimality gap better than $1/\sqrt{N}$. In this paper, we identify conditions analogous to those for RBs under which a better-than-$1/\sqrt{N}$ optimality gap is achievable, and design a policy that attains an $O(1/N)$ optimality gap. Notably, unlike prior approaches based on generalizing priority orderings, our policy is not priority-based but rather is designed to induce locally linear mean-field dynamics.
Problem

Research questions and friction points this paper is trying to address.

Average-reward
Weakly-coupled MDPs
Optimality gap
Restless bandits
Innovation

Methods, ideas, or system contributions that make the work stand out.

Weakly-Coupled MDPs
Average-Reward
Optimality Gap
Mean-Field Dynamics
Restless Bandits
🔎 Similar Papers
No similar papers found.