Breaking the Total Variance Barrier: Sharp Sample Complexity for Linear Heteroscedastic Bandits with Fixed Action Set

📅 2026-07-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses heteroscedastic linear bandits with a fixed action set, where existing methods suffer from dependence on the cumulative variance Λ, which fails to capture the true statistical complexity. The authors propose VAEE, a variance-adaptive algorithm that integrates variance-aware G-optimal design, information-gain-maximizing exploration, and an action elimination mechanism. VAEE is the first to break the √Λ dependence barrier, achieving a simple regret bound that depends nearly only on the harmonic mean of the variances. This bound is nearly optimal for large action sets and exhibits improved dependence on the dimension d in the finite-action setting. Furthermore, the paper establishes a matching information-theoretic lower bound, demonstrating that the attained rate is nearly optimal for fixed action sets.
📝 Abstract
Recent years have witnessed increasing interests in tackling heteroscedastic noise in bandits and reinforcement learning. In these works, the cumulative variance of the noise $Λ= \sum_{t=1}^T σ_t^2$, where $σ_t^2$ is the variance of the noise at round $t$, is used to characterize the statistical complexity of the problem, yielding \emph{simple regret} bounds of order $\tilde{\cal{O}}(d \sqrt{Λ/ T^2})$ for $d$-dimensional linear bandits with heteroscedastic noise. However, with a closer look, $Λ$ remains the same order even if the noise is close to zero at half of the rounds, which indicates that the $Λ$-dependence is not optimal. In this paper, we revisit the stochastic linear bandit problem with heteroscedastic noise, where the action set is prefixed throughout the learning process. We propose a novel variance-adaptive algorithm \texttt{VAEE} (Variance-Aware Exploration with Elimination) for large action set, which actively explores actions that maximizes the information gain among a candidate set of actions that are not eliminated. With the active-exploration strategy, we show that \texttt{VAEE} achieves a \emph{simple regret} with a nearly \emph{harmonic-mean} dependent rate. For finitely many actions, we propose a variance-aware variant of G-optimal design based exploration, which achieves a simple regret with sharper dependence on $d$. We also establish a nearly matching lower bound for the fixed action set setting indicating that \emph{harmonic-mean} dependent rate is unavoidable. To the best of our knowledge, this is the first work that breaks the $\sqrtΛ$ barrier for stochastic linear bandits with heteroscedastic noise.
Problem

Research questions and friction points this paper is trying to address.

heteroscedastic bandits
linear bandits
simple regret
variance dependence
fixed action set
Innovation

Methods, ideas, or system contributions that make the work stand out.

heteroscedastic bandits
variance-adaptive algorithm
harmonic-mean dependence
simple regret
fixed action set
🔎 Similar Papers
2024-10-02International Conference on Machine LearningCitations: 1