Polylogarithmic Nash Regret in Matrix Games with Bandit Feedback

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of minimizing Nash regret under bandit feedback in unknown finite matrix games, with a focus on overcoming the challenge of non-unique equilibria. To this end, it proposes an optimistic payoff-balancing algorithm that constructs reference strategies and applies uncertainty scaling based on estimation accuracy to handle equilibrium multiplicity. Notably, this approach extends polylogarithmic performance guarantees from 2×2 games to arbitrary dimensions. The primary contribution of this work lies in resolving a long-standing open problem in the field by achieving an instance-dependent O(log²T) Nash regret bound. This result outperforms the best known guarantees against adversarial adaptive opponents, thereby establishing a new state-of-the-art for learning in general-sum matrix games under partial information.
📝 Abstract
We study Nash regret minimization in unknown finite matrix games with bandit payoff feedback and observed opponent actions. We develop Optimistic Payoff Balancing (OPB), which achieves instance-dependent $\mathcal{O}(\log^2 T)$ Nash regret against arbitrary adaptive opponents, including games with nonunique equilibria. This resolves the open problem posed by Maiti et al. (2025), extending their polylogarithmic guarantee under bandit feedback from $2\times2$ games to arbitrary finite dimensions. To handle nonunique equilibria, we construct a reference strategy that leaves room for local adjustments. We order independent payoff differences by estimation accuracy and scale these adjustments by uncertainty, allowing the learner to exploit the opponent's imbalance to offset estimation costs. Our result thus shows that observing opponent actions suffices for polylogarithmic Nash regret in general finite matrix games.
Problem

Research questions and friction points this paper is trying to address.

Nash regret minimization
matrix games
bandit feedback
polylogarithmic regret
nonunique equilibria
Innovation

Methods, ideas, or system contributions that make the work stand out.

Nash regret minimization
Optimistic Payoff Balancing
Bandit feedback
Matrix games
Polylogarithmic regret
🔎 Similar Papers
No similar papers found.