On the Complexity of Preference-Based Bandits

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses preference learning under general reward function classes in contextual dueling bandits, overcoming the complexity induced by link function nonlinearity and the limitations of existing methods restricted to linear models. We introduce the locally sensitive Eluder dimension as a novel complexity measure and propose the GINOP algorithm, which integrates the Bradley-Terry model, log-loss confidence sets, and optimistic optimization for efficient exploration. Theoretically, we demonstrate that this framework eliminates the unfavorable dependence on the problem-dependent constant κ, revealing that learning from preference feedback is statistically as efficient as learning from direct observations. Furthermore, we establish a first-order regret bound independent of κ. Empirical evaluations confirm that the proposed algorithm significantly outperforms existing baselines.
📝 Abstract
We study preference-based bandits with general reward function classes, where a learner sequentially selects pairs of arms and observes binary preference feedback governed by the Bradley--Terry model. This setting naturally arises in applications such as recommender systems, tournament ranking, and learning from human feedback, where relative preferences are easier to elicit than absolute rewards. The observation model inherits the logistic bandit challenge of handling the problem-dependent constant $κ$, which accounts for the non-linearity of the link function and can grow arbitrarily large. Moreover, prior work has predominantly focused on linear or kernelized reward models, precluding the use of richer function classes. To address these limitations, we consider general reward function classes and introduce the \emph{locally sensitive eluder dimension}, a novel complexity measure tailored to the logistic structure of preference feedback that yields fine-grained regret guarantees without unfavorable dependence on $κ$. Building on this notion, we propose \textbf{GINOP} (Generic INformative OPtimism), an algorithm that constructs log-loss confidence sets and jointly selects arm pairs to balance optimism and informative exploration. We establish a first-order regret bound that, in contrast with what previous results suggest, demonstrates that learning with preference feedback is as statistically efficient as learning from direct reward observation. Finally, we corroborate our theoretical findings with empirical evaluations against competitive baselines.
Problem

Research questions and friction points this paper is trying to address.

preference-based bandits
Bradley-Terry model
general reward function classes
problem-dependent constant
regret bound
Innovation

Methods, ideas, or system contributions that make the work stand out.

Preference-Based Bandits
Locally Sensitive Eluder Dimension
GINOP
First-Order Regret Bound
Log-Loss Confidence Sets
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.