Saturation-Insensitive Dueling Bandits with General Function Approximation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of feedback information and the deterioration of sample complexity caused by preference saturation in the Bradley-Terry-Luce (BTL) model. To overcome this, we propose the SI-CDB algorithm, which achieves saturation-insensitive reward learning for contextual dueling bandits through a heuristic opponent selection strategy. Furthermore, we introduce a local Eluder dimension framework that elucidates the critical role of dueling regret in determining single-arm performance. Our contributions demonstrate that, under general function approximation and linear reward classes, the proposed approach eliminates the unfavorable 1/σ' factor, recovers near-optimal sample complexity dependencies, and significantly improves reward estimation efficiency.
📝 Abstract
We study contextual dueling bandits with general function approximation under the Bradley-Terry-Luce (BTL) preference model. A key challenge in this setting is the saturation of the preference model: when the current reward model can already distinguish two actions with high confidence, the resulting preference feedback becomes weakly informative, making it difficult to further improve reward estimation. Consequently, existing sample-complexity analyses often depend on the inverse-derivative factor $1 / \sigma'[\Delta_{r^\ast}]$ which can be prohibitively large when the link function $\sigma$ saturates for large reward gaps $\Delta_{r^\ast}$. To address this issue, we introduce `SI-CDB`, an algorithm that selects opponent arms using a carefully designed heuristic for arm selection. This design enables saturation-insensitive reward learning and recovers the near-optimal dependence for linear reward classes, eliminating the unfavorable $1/\sigma'(\cdot)$ factor. The core of our analysis is a localized Eluder dimension framework tailored to dueling bandits with general function approximation. Our theoretical results also explain why two-arm regret analysis is crucial for improving single-arm performance in dueling bandits.
Problem

Research questions and friction points this paper is trying to address.

Contextual Dueling Bandits
General Function Approximation
Preference Model Saturation
Sample Complexity
Bradley-Terry-Luce Model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contextual Dueling Bandits
Saturation-Insensitive Learning
General Function Approximation
Localized Eluder Dimension
Bradley-Terry-Luce Model
🔎 Similar Papers
2024-07-24arXiv.orgCitations: 4
💼 Related Jobs
No related jobs found.
C
Chenggong Zhang
Department of Computer Science, University of California, Los Angeles, CA 90095, USA
X
Xuheng Li
Department of Computer Science, University of California, Los Angeles, CA 90095, USA
Q
Qiwei Di
Department of Computer Science, University of California, Los Angeles, CA 90095, USA
Weitong Zhang
Weitong Zhang
Assistant Professor, SDSS, UNC Chapel Hill
Reinforcement LearningOptimizationAI4Science
Quanquan Gu
Quanquan Gu
Associate Professor of Computer Science, UCLA
AGILarge Language ModelsReinforcement LearningNonconvex Optimization