Score
Designs and analyzes learning policies and algorithms that minimize policy regret (rp-regret) in repeated interactions by comparing cumulative realized utility to the best-in-hindsight history-dependent policy rather than to a single fixed action. Work includes deriving minimax-optimal guarantees and regret-optimal policy constructions that account for opponents who adapt to play histories and that enable equilibrium learning in repeated games when policy regret is driven to zero.
This work addresses the limitation of traditional external regret in repeated games, which fails to capture scenarios where opponents adaptively adjust their strategies based on history. The authors propose a novel regret measure, RP-Regret, that quantifies the gap between a player’s actual cumulative utility against an adaptive opponent and the utility achievable by the best ex post response. This measure enables a stronger performance benchmark under weaker assumptions about opponent behavior and can lead to more efficient equilibria. To handle the non-convexity inherent in RP-Regret minimization, the paper integrates techniques from non-convex optimization, linearized surrogate objectives, optimization oracles, and direct optimization under the assumption of slowly changing opponents. Theoretically, RP-Regret is shown to grow sublinearly under suitable conditions, and experiments demonstrate that the approach effectively learns high-payoff cooperative equilibria in games such as Stag-Hunt.
This work addresses decentralized learning in multi-agent reinforcement learning under both general games and Markov games. It is the first to achieve sublinear swap regret minimization—guaranteeing convergence of policy sequences to correlated equilibria—under fully decentralized, communication-free, and identical-policy-update settings. Methodologically, it introduces a weighted regret framework with path-length-adaptive weighting, integrating online policy optimization and decentralized learning mechanisms. The approach overcomes theoretical bottlenecks in adversarial Markov games for regret minimization, remaining statistically and computationally feasible. Theoretically, it achieves an $O(sqrt{T})$ swap regret bound—strictly improving upon prior results. Crucially, the algorithm requires no inter-agent communication, ensuring strong practicality and scalability across large-scale multi-agent systems.
This paper investigates the design of optimal strategies for a leader facing a no-regret learner in repeated Stackelberg games. We employ game-theoretic modeling, Stackelberg equilibrium analysis, and counterfactual utility bounding to characterize the leader’s guaranteed utility. Our contributions are threefold: (i) We derive the first tight upper bound on the leader’s achievable utility and identify precise conditions under which this bound can be exceeded; (ii) We construct an optimal leader strategy for the three-action setting; (iii) We prove that in the two-action case, the Stackelberg equilibrium utility is theoretically optimal, whereas with multi-action mean-based no-regret learners, the leader can strictly surpass this benchmark; conversely, under no-swap regret constraints, the Stackelberg utility constitutes an unattainable upper bound. Collectively, our results unify the understanding of utility limits and attainability for leaders across distinct regret models—no-regret, mean-based no-regret, and no-swap regret—thereby clarifying fundamental trade-offs in strategic learning interactions.
Large language model (LLM) agents lack quantitative, rational evaluation frameworks in interactive decision-making settings—such as online learning and multi-agent games—hindering rigorous assessment of their strategic competence. Method: This work introduces “regret” as a foundational metric to systematically characterize LLMs’ decision-making boundaries. We propose an unsupervised “regret loss” training objective—requiring no action-level supervision—and ground it theoretically via generalization bounds and convergence analysis. Integrating game-theoretic modeling, statistical learning theory, and optimization analysis, we conduct experiments on nonstationary online learning and repeated games using GPT-4. Contribution/Results: Empirical results reveal substantial cumulative regret even in simple games, exposing critical limitations in current LLM rationality. Regret loss training significantly reduces regret and accelerates Nash equilibrium emergence. This establishes a novel paradigm for rational modeling and alignment of LLMs, bridging theoretical guarantees with practical agent behavior.
This paper addresses strategy equilibrium in repeated games under dynamic environments, challenging the conventional assumption of static optimal actions. Method: We introduce the novel concept of “dynamic benchmark consistency,” which permits a bounded number of action switches to approximate an optimal dynamic action sequence. Integrating empirical distribution analysis, online learning theory, and dynamic programming–based constraint modeling, we formulate and analyze adaptive strategies under evolving benchmarks. Contribution/Results: We provide the first rigorous proof that dynamically benchmark-consistent strategies generate precisely the same Nash-type equilibrium set—asymptotically—as classical no-regret strategies. This establishes an intrinsic unification between stringent individual regret constraints and collective coordination equilibria. Moreover, it demonstrates that independently adaptive algorithms can spontaneously achieve strong coordination in large-horizon settings, without explicit communication or centralized control. Our results furnish a new theoretical foundation for strategy design and mechanism interpretation in dynamic games.
This study addresses how intelligent agents can achieve rational decision-making in unknown, non-stationary, and even adversarial environments, and investigates the emergence and stability of equilibria in multi-agent interactions. The work proposes a unified analytical framework that integrates regularized learning strategies, adversarial multi-armed bandit models, and fictitious play dynamics, accommodating both oracle information and bandit feedback settings. Key contributions include deriving optimal regret bounds for single-agent learning in adversarial environments, establishing ergodic convergence to Nash equilibria in zero-sum games, and formulating a folk-theorem-like correspondence between attractors of regularized learning dynamics and Nash equilibria, thereby revealing a fundamental alignment between strategic stability and adaptive learning behavior.
This study investigates how to effectively leverage potentially imperfect machine learning advice to enhance strategic performance in repeated games against no-regret learners. By introducing a pseudometric that quantifies the quality of advice, the work systematically analyzes two types of recommendations—simulators and payoff matrix predictors—in two-player repeated settings. The theoretical results demonstrate that when advice comes with correctness guarantees, an approximate Stackelberg strategy can be computed efficiently; without such guarantees, it is impossible to simultaneously achieve near-optimality and no regret, though weakly dominant utilities can still be attained within certain (coarse) correlated equilibria. This work establishes the first theoretical limits on the use of unverified advice and reveals that high-quality advice can substantially reduce interaction complexity.
This work studies efficient coordination under asymmetric information in repeated online assisted games with a shared reward, where the informed player (human) observes a hidden state while the uninformed assistant only sees the human’s actions. The paper introduces, for the first time, a provably efficient decentralized learning algorithm that leverages the novel notion of “assistance regret.” It establishes that achieving a $(1 - 1/e)$ approximation factor is computationally infeasible in this setting. By combining any no-regret learner with a shared random seed, the proposed method attains an optimal $\widetilde{O}(T^{1/2})$ regret bound in a pseudo-decentralized framework and achieves a $(1 - 1/e)$-approximate assistance regret bound of $\widetilde{O}(T^{3/4})$.
When data are insufficient to learn policies with low regret or significantly better performance than a baseline, how can we characterize the intrinsic difficulty of policy learning? This work proposes a unified framework to systematically study three fundamental problems: optimal policy learning, improved policy learning, and policy existence verification. Through theoretical analysis, problem reductions, and sample complexity comparisons, the paper establishes a strict or partially strict hierarchy of difficulty among these tasks: optimal policy learning is provably harder than improved policy learning, and under natural conditions, a sublinear polynomial complexity gap separates improved policy learning from existence verification. Notably, this study formalizes the policy existence problem for the first time and reveals that even when constructing an improved policy is infeasible, efficiently determining its existence may still be possible.
This work investigates how to train models in an unsupervised setting to achieve no-regret and swap-regret properties from game theory, thereby inducing equilibrium behavior. The authors introduce a novel approach by directly formulating external regret and swap regret as differentiable loss functions and embedding them within a single-layer self-attention architecture. This design ensures that forward propagation is equivalent to smoothed fictitious play and its swap variant. The framework naturally recovers the update rules of classical online learning algorithms such as those of Blum and Mansour, accurately simulating regret dynamics without explicit supervision. Consequently, it guarantees convergence to coarse correlated equilibria and even correlated equilibria, revealing a profound connection between attention mechanisms and game-theoretic equilibrium concepts.