🤖 AI Summary
This study addresses best arm identification in stochastic multi-armed bandits under adversarial mean shifts, demonstrating that conventional Generalized Likelihood Ratio Tests (GLRT) and Track-and-Stop algorithms fail within the Shifting Means model. To overcome this limitation, this work proposes a novel Importance Sampling for Means (ISM) framework that leverages importance-weighted sampling combined with sub-Gaussian analysis to minimize sample complexity in the fixed-confidence setting. The authors rigorously prove that the ISM algorithm satisfies δ-correctness, establishing a sample complexity upper bound of K(σ²+U²)Δ_min⁻²ln(1/δ) alongside a matching lower bound. Empirical evaluations further validate the practical effectiveness of the proposed approach.
📝 Abstract
We study the best arm identification problem in a stochastic environment with a novel form of adversarial perturbations, which we coin Shifting Means. While classically the mean rewards of the $K$ arms are stable in time, in Shifting Means only the gaps $\boldsymbolΔ$ between mean rewards are stable, while their common shift may be determined adversarially in each round. The objective of the learner is to identify the best arm with high probability while minimizing sample complexity (the fixed confidence setting). Handling shifts requires new tools: we show that algorithms employing a Generalized Likelihood Ratio Test (GLRT) stopping rule, including the popular Track-and-Stop, fail under time-varying shifts. Instead, we propose Importance Weights for Shifting Means ($\mathsf{ISM}$). Assuming means bounded by $U$ and $σ^2$-sub-Gaussian rewards, we show $\mathsf{ISM}$ to be $δ$-correct and to enjoy a sample complexity bound of order $K (σ^2 + U^2) Δ_{\min}^{-2} \ln \frac{1}δ$. We also present a matching (up to constant factors) worst-case lower bound and evaluate our results empirically.