Learning in Continuous Games from Pairwise Preference Feedback

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the non-identifiability of traditional external regret and coarse correlated equilibria (CCE) in continuous games where only ordinal preference feedback is available. To overcome this limitation, it proposes a first-order ordinal framework based on normalized preference directions and introduces an ordinal directional regret benchmark that eliminates the need to reconstruct cardinal utilities. Furthermore, the work designs block-normalized pseudo-gradient dynamics coupled with a single-comparison estimator to enable efficient online learning. Theoretically, this approach achieves sublinear ordinal regret bounds and establishes almost sure convergence to the Nash set in potential games. Overall, this research presents a novel paradigm for equilibrium computation in multi-agent games under purely ordinal feedback.
📝 Abstract
We study learning in continuous games when players receive only pairwise preference feedback, revealing which of two actions is preferred but neither payoff values nor preference magnitudes. We first show that standard external regret and coarse correlated equilibria (CCE) are not identifiable from this ordinal information: the same sequence of play can incur zero and linear regret in two ordinally equivalent games, while the distributions that remain CCE across all cardinal representations consistent with the same preferences are exactly those supported on pure Nash equilibria. Motivated by this gap, we develop a first-order ordinal theory based on normalized unilateral preference directions, introducing an ordinal directional regret benchmark and corresponding equilibrium notions. We show that block-normalized pseudogradient dynamics achieve sublinear ordinal regret and, under additional structure, Nash-convergence guarantees. We then use a single-comparison estimator to implement these dynamics from finite pairwise comparisons. With one comparison per player and round, the resulting algorithm achieves sublinear finite-resolution ordinal regret against arbitrary opponent behavior and, in ordinal potential games, almost-sure last-iterate convergence to the Nash set. Our results provide regret, dynamics, and equilibrium guarantees directly from preference feedback without reconstructing cardinal utilities.
Problem

Research questions and friction points this paper is trying to address.

continuous games
pairwise preference feedback
ordinal regret
coarse correlated equilibria
Nash equilibrium
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pairwise Preference Feedback
Ordinal Directional Regret
Continuous Games
Pseudogradient Dynamics
Single-Comparison Estimator
🔎 Similar Papers
No similar papers found.