🤖 AI Summary
This study addresses the reliance on discretization and poor sample efficiency of existing algorithms in continuous-action games. It proposes a scalable policy gradient method that, for the first time, integrates magnetic mirror descent with Gaussian mixture model reparameterization. Through self-play reinforcement learning, this approach effectively approximates Nash equilibria in continuous and mixed-action sequential games, even under gradient failure conditions. The proposed method substantially enhances solution capabilities in continuous action spaces. Compared to neural fictitious self-play and Policy-Space Response Oracles (PSRO), it achieves a 3.5- to 5.5-fold improvement in sample efficiency. Furthermore, its performance in Texas Hold’em poker is comparable to that of Slumbot, demonstrating strong practical efficacy in complex game-theoretic settings.
📝 Abstract
Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable policy-gradient algorithm for large sequential games with continuous or mixed discrete and continuous actions. It combines magnetic mirror descent with a mixture of Gaussians reparametrization, trained via self-play. We show that it approximates equilibrium in games where gradient descent fails. In sequential games, it outperforms neural fictitious self-play and matches or outperforms the final strategies of policy space response oracles with 3.5--5.5$\times$ fewer samples. In heads-up no-limit Texas hold'em, it performs on par with Slumbot.