🤖 AI Summary
This study addresses the challenges of intractable minimax optimization and the absence of finite-sample guarantees in continuous-action zero-sum Markov games, where existing theoretical results are largely confined to linear-quadratic (LQ) settings. To tackle non-LQ games, this work proposes a convex-concave neural network function approximation class that ensures the existence of saddle points, and designs an online fitted Q-iteration algorithm to enable efficient policy learning. The primary contribution lies in establishing the first finite-sample convergence guarantees for non-LQ continuous state-action zero-sum Markov games. Empirical evaluations and open-sourced code further validate the effectiveness of the proposed approach.
📝 Abstract
Zero-sum Markov games arise in a wide variety of sequential decision-making problems such as adversarial learning and planning against modeled uncertainties. However, prior work on finite-sample guarantees on the learned state-action value function ($Q$-function) for zero-sum Markov games is largely restricted to finite-action settings, or to continuous-action games in which agents have linear dynamics and quadratic rewards (i.e. linear-quadratic, or LQ, games). A central challenge in extending such guarantees to more general continuous-action zero-sum Markov games is that the associated Bellman operator involves a minimax problem that need not admit a tractable saddle-point solution. To this end, we first introduce a class of neural network function approximators for the $Q$-function that is convex-concave in the players'actions, which guarantees that this minimax problem admits a pure-strategy saddle point. We then study an online variant of fitted $Q$-iteration employing this function class and establish, to the best of our knowledge, the first finite-sample guarantees for non-LQ zero-sum Markov games with continuous states and actions. Our code can be found at https://github.com/CLeARoboticsLab/QCanPlayThatGame.