Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance on discretization and poor sample efficiency of existing algorithms in continuous-action games. It proposes a scalable policy gradient method that, for the first time, integrates magnetic mirror descent with Gaussian mixture model reparameterization. Through self-play reinforcement learning, this approach effectively approximates Nash equilibria in continuous and mixed-action sequential games, even under gradient failure conditions. The proposed method substantially enhances solution capabilities in continuous action spaces. Compared to neural fictitious self-play and Policy-Space Response Oracles (PSRO), it achieves a 3.5- to 5.5-fold improvement in sample efficiency. Furthermore, its performance in Texas Hold’em poker is comparable to that of Slumbot, demonstrating strong practical efficacy in complex game-theoretic settings.
📝 Abstract
Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable policy-gradient algorithm for large sequential games with continuous or mixed discrete and continuous actions. It combines magnetic mirror descent with a mixture of Gaussians reparametrization, trained via self-play. We show that it approximates equilibrium in games where gradient descent fails. In sequential games, it outperforms neural fictitious self-play and matches or outperforms the final strategies of policy space response oracles with 3.5--5.5$\times$ fewer samples. In heads-up no-limit Texas hold'em, it performs on par with Slumbot.
Problem

Research questions and friction points this paper is trying to address.

continuous actions
sequential games
policy gradient
sample efficiency
equilibrium approximation
Innovation

Methods, ideas, or system contributions that make the work stand out.

policy gradient
mixture of Gaussians
magnetic mirror descent
self-play
continuous actions