🤖 AI Summary
This study addresses the challenge of last-iterate convergence in unknown two-player zero-sum matrix games under bandit feedback with observable opponent actions, proposing a computationally efficient algorithm. Methodologically, the approach combines adaptive averaging with a modified exponential weights update strategy, and employs potential function analysis for theoretical characterization. The results demonstrate that the proposed algorithm achieves an O(√(d/t)) duality gap with high probability, while requiring only O(d) time and space complexity per round. Notably, its dependence on the dimensionality improves upon existing work by a factor of d^{3/2}, matching the information-theoretic lower bound and thereby establishing minimax optimality.
📝 Abstract
We study last-iterate convergence in unknown two-player zero-sum matrix games with bandit payoff feedback and observed opponent actions. For games with $d$ actions per player, we develop an algorithm achieving a duality gap of $\widetilde{\mathcal{O}}(\sqrt{d/t})$ with high probability, simultaneously at every round $t$. This improves the dimension dependence of the best previously known guarantee by a factor of $d^{3/2}$. The rate matches a standard bandit lower bound, establishing minimax optimality in both the number of actions and the number of rounds, up to logarithmic factors. The algorithm is computationally efficient, requiring only $\mathcal{O}(d)$ time and memory per round. Our technical contribution is a joint design of adaptive averaging and corrected exponential weights that absorbs estimation variance, together with a potential argument that bounds phase durations.