🤖 AI Summary
This study addresses the non-identifiability of traditional external regret and coarse correlated equilibria (CCE) in continuous games where only ordinal preference feedback is available. To overcome this limitation, it proposes a first-order ordinal framework based on normalized preference directions and introduces an ordinal directional regret benchmark that eliminates the need to reconstruct cardinal utilities. Furthermore, the work designs block-normalized pseudo-gradient dynamics coupled with a single-comparison estimator to enable efficient online learning. Theoretically, this approach achieves sublinear ordinal regret bounds and establishes almost sure convergence to the Nash set in potential games. Overall, this research presents a novel paradigm for equilibrium computation in multi-agent games under purely ordinal feedback.
📝 Abstract
We study learning in continuous games when players receive only pairwise preference feedback, revealing which of two actions is preferred but neither payoff values nor preference magnitudes. We first show that standard external regret and coarse correlated equilibria (CCE) are not identifiable from this ordinal information: the same sequence of play can incur zero and linear regret in two ordinally equivalent games, while the distributions that remain CCE across all cardinal representations consistent with the same preferences are exactly those supported on pure Nash equilibria. Motivated by this gap, we develop a first-order ordinal theory based on normalized unilateral preference directions, introducing an ordinal directional regret benchmark and corresponding equilibrium notions. We show that block-normalized pseudogradient dynamics achieve sublinear ordinal regret and, under additional structure, Nash-convergence guarantees. We then use a single-comparison estimator to implement these dynamics from finite pairwise comparisons. With one comparison per player and round, the resulting algorithm achieves sublinear finite-resolution ordinal regret against arbitrary opponent behavior and, in ordinal potential games, almost-sure last-iterate convergence to the Nash set. Our results provide regret, dynamics, and equilibrium guarantees directly from preference feedback without reconstructing cardinal utilities.