Neural Exploitation and Exploration of Contextual Bandits

📅 2023-05-05
🏛️ arXiv.org
📈 Citations: 9
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the exploration-exploitation trade-off in contextual multi-armed bandits. We propose EE-Net, a dual-neural-network architecture: one network models the reward function for efficient exploitation, while the other directly learns instance-dependent exploration gains—bypassing conventional statistical confidence bounds—to enable adaptive exploration. To our knowledge, this is the first work to explicitly model exploration gains using neural networks. Theoretical analysis establishes an instance-dependent regret upper bound of $ ilde{O}(sqrt{T})$. Empirical evaluation on multiple real-world datasets demonstrates that EE-Net significantly outperforms both linear and state-of-the-art neural contextual bandit baselines, validating its modeling flexibility and generalization capability.
📝 Abstract
In this paper, we study utilizing neural networks for the exploitation and exploration of contextual multi-armed bandits. Contextual multi-armed bandits have been studied for decades with various applications. To solve the exploitation-exploration trade-off in bandits, there are three main techniques: epsilon-greedy, Thompson Sampling (TS), and Upper Confidence Bound (UCB). In recent literature, a series of neural bandit algorithms have been proposed to adapt to the non-linear reward function, combined with TS or UCB strategies for exploration. In this paper, instead of calculating a large-deviation based statistical bound for exploration like previous methods, we propose, ``EE-Net,'' a novel neural-based exploitation and exploration strategy. In addition to using a neural network (Exploitation network) to learn the reward function, EE-Net uses another neural network (Exploration network) to adaptively learn the potential gains compared to the currently estimated reward for exploration. We provide an instance-based $widetilde{mathcal{O}}(sqrt{T})$ regret upper bound for EE-Net and show that EE-Net outperforms related linear and neural contextual bandit baselines on real-world datasets.
Problem

Research questions and friction points this paper is trying to address.

Proposes EE-Net for neural contextual bandit exploitation and exploration
Uses separate neural networks to adaptively learn reward and exploration gains
Achieves sublinear regret and outperforms existing baselines on real datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neural network learns reward function for exploitation
Separate neural network adaptively learns exploration potential
Achieves sublinear regret bound without statistical deviation bounds
🔎 Similar Papers
2024-07-24arXiv.orgCitations: 4