Which Nash Equilibrium? Solver-Dependent Selection on Zero-Sum Nash Polytopes

📅 2026-06-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether different solvers systematically favor particular equilibria when multiple Nash equilibria exist in two-player zero-sum games. The authors construct benchmark games with analytically tractable Nash equilibrium sets and systematically evaluate the equilibrium selection behavior of algorithms including R-NaD, mirror descent, CFR, CFR+, and fictitious play. Their findings reveal that equilibrium selection is primarily governed by algorithmic type rather than random initialization: R-NaD consistently converges to the maximum-entropy equilibrium, whereas regret-minimization methods such as CFR+ tend to select low-entropy boundary equilibria. Further experiments demonstrate that the maximum-entropy equilibrium exhibits greater robustness against suboptimal opponents in Kuhn poker, highlighting the advantage of regularized last-iterate methods in achieving higher-quality equilibria.
📝 Abstract
Many two-player zero-sum games admit not a unique Nash equilibrium but a convex set of them: a polytope of profiles that all share the minimax value V* yet prescribe different behaviour. Standard solvers each converge to some equilibrium and are treated as interchangeable. We ask whether they instead select different members of the Nash set, systematically as a function of the algorithm rather than the seed. Using a tabular, exactly solvable testbed of six games with analytically known Nash sets -- including a two-dimensional Nash polytope and Kuhn poker -- we find that (i) selection is determined by the algorithm, not the seed, but families differ only on asymmetric Nash sets; (ii) regularized last-iterate methods (R-NaD, magnetic mirror descent) select the maximum-entropy member, the information projection of their uniform reference onto the Nash set -- exactly on the 2-D polytope and at 99.7% of maximum entropy in Kuhn -- while regret-averaging methods (CFR, CFR+, fictitious play) drift to a lower-entropy face; we confirm this on a randomized 180-game ensemble, where R-NaD attains the maximum-entropy member in 100% of converged games while CFR+ sits strictly below it in 94% (paired Wilcoxon p < 10^-27); (iii) the selected member has downstream consequences against sub-optimal opponents that scale with sequential/hidden-information structure but stay bounded -- in Kuhn the max-entropy member is a strictly better hedge, whereas on the matrix games the members differ without either dominating. We also report two negative results correcting common intuitions: removing CFR's positive-orthant (max(R,0)) projection does not eliminate boundary drift; and R-NaD's selection is anchor-following, not initialization-independent. We state the maximum-entropy / I-projection characterization as a strongly data-supported conjecture, checked throughout against analytic ground truth.
Problem

Research questions and friction points this paper is trying to address.

Nash equilibrium selection
zero-sum games
solver-dependent behavior
Nash polytope
equilibrium multiplicity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Nash equilibrium selection
maximum entropy
zero-sum games
solver-dependent behavior
information projection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Luis Leal