AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play

πŸ“… 2026-05-09
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of enhancing payoff by effectively exploiting suboptimal opponent behaviors while maintaining robustness against Nash equilibrium strategies. To this end, the authors propose a novel approach that integrates a hierarchical Transformer encoder with adversarial reinforcement learning. By modeling historical interaction data and leveraging a diverse pool of exploitable opponents, the agent dynamically adapts its policy to target specific weaknesses in adversaries’ strategies. This method uniquely combines hand-history reasoning with adversarial reinforcement learning, achieving state-of-the-art performance on standard imperfect-information game benchmarks. It demonstrates significant improvements over existing techniques by efficiently exploiting both in-distribution and out-of-distribution suboptimal opponents without compromising robustness to equilibrium play.
πŸ“ Abstract
Poker is an imperfect information game that has served as a long-standing benchmark for decision-making under uncertainty. To maximize utility beyond the Nash equilibrium, an agent can deviate from Nash-equilibrium policies to exploit suboptimal play. We introduce AlphaExploitem, which extends the competitive RL poker agent AlphaHoldem by using a hierarchical transformer encoder that enables reasoning over previously played hands and modifying the training procedure with the inclusion of a diverse pool of exploitable opponents to facilitate learning to exploit. We train and evaluate AlphaExploitem on two standard benchmarks for imperfect-information games. Empirically, AlphaExploitem successfully exploits weak play by both in- and out-of-distribution opponents, without losing performance against NE opponents.
Problem

Research questions and friction points this paper is trying to address.

poker
Nash equilibrium
exploitation
imperfect information games
suboptimal play
Innovation

Methods, ideas, or system contributions that make the work stand out.

exploitability
imperfect information games
hierarchical transformer
Nash equilibrium deviation
opponent modeling
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
V
Vlad Murgoci
Delft University of Technology
M
Matthijs Spaan
Department of Intelligent Systems, Delft University of Technology
Yaniv Oren
Yaniv Oren
PhD candidate, Delft University of Technology
Reinforcement Learning