Pessimistic Minimax Learning for Public-Private Information Games under Unilateral Coverage

📅 2026-10-04
📈 Citations: 2
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of a theoretical framework for offline equilibrium learning in zero-sum games with public and private information. To this end, it introduces the concept of one-sided coverability and constructs a Pessimistic Policy Mirror Descent (PPA-PMD) framework based on pessimistic estimation and function approximation to achieve no-regret policy updates. This work establishes the first theory for offline equilibrium learning under asymmetric information constraints, revealing how such asymmetry influences data coverage mechanisms. Furthermore, it demonstrates that the proposed method attains an exploitability rate of $\tilde{O}(1/\sqrt{n})$, matching the convergence speed of fully observable games and thereby unifying the theoretical bounds.
📝 Abstract
We study offline learning in two-player zero-sum contextual games with public and private information, motivated by strategic settings such as auctions and negotiations with private valuations. We introduce unilateral prescriptive concentrability and show that asymmetric information can change offline coverage through its effect on equilibrium behavior. For finite state-action spaces, we develop a pessimistic algorithm with an $\tilde{O}(1/\sqrt{n})$ exploitability rate, matching the standard sample-size dependence for fully observed minimax games. We further develop a pessimistic policy mirror descent framework, PPA-PMD, for general function approximation and obtain a unified $\tilde{O}(1/\sqrt{n} + 1/\sqrt{T})$ exploitability rate with no-regret actor updates. Together, these results provide the first theoretical framework for offline equilibrium learning under public-private information constraints.
Problem

Research questions and friction points this paper is trying to address.

offline learning
zero-sum contextual games
public-private information
asymmetric information
equilibrium learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Offline Learning
Pessimistic Minimax
Public-Private Information Games
Policy Mirror Descent
Unilateral Prescriptive Concentrability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shuze Daniel Liu
Massachusetts Institute of Technology, Purdue University
Claire Chen
Claire Chen
PhD student, Stanford University
contact-rich manipulationrobot learningmulti-modal sensing
Jiuqi Wang
Jiuqi Wang
Ph.D. Student, University of Virginia
reinforcement learningdeep learning
D
David Simchi-Levi
Massachusetts Institute of Technology, Purdue University