projected exploitability descent

Designs and implements iterative optimization methods that minimize a projected proxy of exploitability for game-playing strategies by computing subgradients of a max-of-linear exploitability surrogate and projecting those subgradients or update steps onto the sequence-form strategy polytope. Builds and analyzes algorithms and their convergence behavior to refine strategies toward Nash equilibria via projected subgradient exploitability minimization.

projectedexploitabilitydescent

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of balancing scalability and convergence in solving Nash equilibria for multiplayer imperfect-information games. The authors propose the Projected Exploitability Descent (PED) algorithm, which, for the first time, expresses the generalized exploitability function as a sum of maxima of linear functions and optimizes it via projected subgradient descent over the sequence-form strategy space, ensuring stable convergence. They further introduce FP-PED, a hybrid approach that combines fictitious play (FP) for warm-start initialization with PED for refined optimization. Experiments demonstrate that PED exhibits nearly monotonic improvement on a three-player variant of Kuhn poker, while FP-PED substantially enhances long-term convergence performance, overcoming the scalability limitations of exact methods when the deck size exceeds four.

computational performanceimperfect-information gamesmultiplayer games

This work addresses the instability of traditional subgame solving in imperfect-information games, where reliance on Nash equilibrium often yields strategies with poor robustness in the full game. To overcome this limitation, the paper proposes using sequential equilibrium—a refinement of Nash equilibrium—as the solution concept for gadget games. By integrating sequence-form linear programming with an enhanced counterfactual regret minimization (CFR) algorithm, the method efficiently converges to this refined equilibrium with only minor additional computational overhead. Experimental results across multiple standard benchmark games demonstrate that the resulting strategies exhibit significantly improved consistency and reduced exploitability—by over 50% compared to unrefined Nash equilibrium strategies—thereby achieving markedly better global performance and robustness.

equilibrium refinementgadget gamesimperfect-information games

This work addresses the open problem posed by Syrgkanis et al. concerning last-iterate convergence of Optimistic Multiplicative Weights Update (OMWU) in constrained convex-concave minimax optimization—particularly relevant to zero-sum games and GANs. We establish, for the first time, global convergence of OMWU to exact saddle points under general convex constraints, without requiring unconstrained domains or strong regularization assumptions. Our analysis introduces a novel framework combining monotonic KL-divergence descent with local contraction mapping properties, integrating fixed-point theory and contraction mapping techniques. This approach overcomes key limitations of prior analyses reliant on either unbounded domains or stringent regularization. The result provides the first rigorous theoretical guarantee for OMWU’s last-iterate convergence in constrained saddle-point optimization and furnishes a principled foundation for termination criteria in practical iterative implementations.

Analyzing Optimistic Multiplicative-Weights Update method for game theoryEstablishing convergence guarantees for saddle point problems with constraintsExtending last-iterate convergence to constrained min-max optimization problems

This paper investigates the convergence of best-response (BR) dynamics in simultaneous-move convex quadratic games over lattices, addressing challenging settings with nonlinear objective functions and unbounded feasible sets. We establish a global convergence criterion based on the singular values of the interaction matrix: BR iterations remain globally bounded if all singular values are less than one; divergence occurs for infinitely many initial points if any singular value exceeds one; and almost-everywhere divergence arises when all singular values exceed one. This yields the first tight singular-value condition guaranteeing BR non-divergence. Furthermore, we introduce the notion of “traps”—finite subgames that confine divergent trajectories—and construct mixed Nash equilibria thereof as relaxation solutions to the original problem. Our results provide a unified spectral characterization of BR convergence and divergence, integrating convex optimization, game theory, and singular value analysis, thereby establishing a theoretical foundation for algorithmic reliability in discrete nonlinear games.

Analyzing convergence of best-response algorithms in lattice convex-quadratic gamesEstablishing conditions for non-divergence based on interaction matrix singular valuesProving divergence when singular values exceed 1 from most initial points

The equilibrium properties of obvious strategy profiles in games with many players

Oct 29, 2024
EC
Enxian Chen
🏛️ Nankai University | Capital University of Economics and Business | Chinese University of Hong Kong

This paper investigates the equilibrium properties of *obvious strategy profiles* in large-scale finite games. Addressing the existence and implementability of approximate symmetric equilibria as the number of players tends to infinity, we propose a fully decentralized, coordination-free constructive method. Under continuity and asymptotic regularity assumptions, we prove that the empirical strategy distributions induced by obvious strategy profiles converge weakly to symmetric approximate Nash equilibria. Moreover, their random pure-strategy realizations constitute pure-strategy approximate Nash equilibria with probability approaching one. This work establishes, for the first time, *dual convergence*: (i) weak convergence of strategy distributions and (ii) high-probability convergence of pure-strategy realizations. The resulting framework yields scalable, robust, and asymptotically optimal equilibrium solutions for large games—circumventing both explicit coordination mechanisms and prohibitive computational complexity inherent in traditional approaches.

Analyzing equilibrium properties in large finite-player gamesProviding easily implemented solutions without coordination issuesStudying convergence of approximate symmetric equilibria asymptotically

Latest Papers

What's happening recently
View more

This work addresses the slow convergence and reliance on strategy space discretization inherent in classical fictitious play algorithms for zero-sum games by proposing Almost Greedy Fictitious Play. The method operates directly in continuous strategy spaces, eschewing discretization, and performs an approximately greedy line search along the segment between the current empirical mixed strategy and its best response to enable more natural strategy updates. Built upon the fictitious play framework and integrating tools from continuous optimization and duality gap analysis, the algorithm achieves instance-dependent theoretical convergence guarantees. It attains a convergence rate of O(1/T) in terms of the duality gap, matching that of continuous fictitious play. Empirical results demonstrate the algorithm’s effectiveness and superiority over existing approaches.

convergence rateduality gapFictitious Play

This study investigates the exploitability of Follow-the-Regularized-Leader (FTRL) learners with fixed step sizes in two-player zero-sum games when facing oracle optimizers. By integrating game theory, online learning theory, and stochastic game models, and distinguishing between fixed and alternating optimizer settings, the work establishes— for the first time—that exploitability is an inherent property of the FTRL family. The core contributions include a geometric dichotomy based on the steepness of the regularizer, a sensitivity metric quantifying vulnerability to strategic manipulation, and theoretical guarantees showing a lower bound of Ω(N/η) on exploitability under fixed optimizers, as well as a high-probability surplus of Ω(ηT/poly(n,m)) in cumulative payoff under alternating optimizers.

exploitabilityFTRL dynamicsregularizers

This work addresses the computation of Nash equilibria in two-player constrained games under asymmetric information, where one player lacks full knowledge of the other’s objective and constraints and can only interact via a best-response mapping. The authors propose an iterative algorithm that combines projected gradient descent with best-response updates, requiring no complete disclosure of either player’s optimization problem and applicable to settings with decoupled feasible sets. Under standard regularity conditions, they establish—for the first time—the global linear convergence of the algorithm in such an asymmetric regime relying solely on best-response access. Moreover, when the best-response is subject to uniformly bounded errors of magnitude ε, the iterates converge to an O(ε)-neighborhood of the equilibrium, with an explicit error bound provided. Numerical experiments corroborate the theoretical convergence rates and sensitivity to response inaccuracies.

asymmetric informationbest-response mapconstrained games

Understanding Optimal Portfolios of Strategies for Solving Two-player Zero-sum Games

Nov 23, 2025
KD
Karolina Drabent
🏛️ Czech Technical University in Prague

This work addresses the problem of constructing compact optimal strategy profiles—comprising a small set of representative strategies—that efficiently approximate the opponent’s strategy space in large two-player zero-sum games, avoiding domain-specific heuristics or methods lacking theoretical guarantees. We establish the first formal theoretical framework for this problem, prove its NP-hardness, and demonstrate that common heuristics—including uniform sampling and support-set expansion—can be severely suboptimal under specific game structures. Our method introduces an evaluation and analysis framework grounded in Nash equilibrium support verification and incremental construction, integrating game-theoretic analysis, computational complexity proofs, and empirical comparisons to derive tight theoretical bounds on strategy profile quality. To foster reproducibility and future research, we release open-source code and standardized benchmark datasets.

Establishing formal foundation for portfolio-based strategy approximation in gamesEvaluating heuristic performance variability across different game typesProving NP-hard complexity of finding optimal strategy portfolios

This work investigates the time complexity of the Follow the Regularized Leader (FTRL) algorithm in converging to Nash equilibria in potential games. By constructing explicit instances, it establishes—for the first time—an exponential lower bound on the convergence time of FTRL in two-player potential games under any permutation-invariant regularizer, and a doubly exponential lower bound in the multi-player setting. The study further reveals that the FTRL dynamics admit a potential function structure and demonstrates their equivalence to mirror descent and fictitious play under appropriate conditions. These results imply that FTRL and its variants, such as multiplicative weights update, require exponential time to converge, whereas lazy alternating no-regret dynamics achieve an upper bound of $\exp(O(1/\varepsilon^2))$, matching the lower bounds up to exponential order.

convergence timeexponential lower boundsFTRL

Hot Scholars

YC

Yixin Cao

Hong Kong Polytechnic University
algorithmsgraph classescombinatorial optimization
EC

Elvan Ceyhan

Auburn University
StatisticsData ScienceProbability
LM

Luke Marris

Research Engineer at DeepMind, PhD from University College London
Machine LearningGame TheoryReinforcement LearningMulti-Agent
ZZ

Zhongyi Zhang

Huazhong University of Science and Technology
LD

Luc Devroye

McGill University
Probabilistic analysis of algorithms