🤖 AI Summary
This work addresses the challenge of accurately identifying the arm with the highest true reward in multi-armed bandit settings, a task where conventional methods often fall short. The authors propose a novel decision-making framework based on a domination score criterion, integrating a joint mixing-and-reuse mechanism with a doubly robust estimator to ensure simultaneous convergence of the empirical distribution functions across all arms. The resulting elimination algorithm achieves near-optimal sample complexity while providing rigorous theoretical guarantees for correctly identifying the true dominant arm. Empirical evaluations demonstrate that the proposed method significantly outperforms existing baselines in both accuracy and efficiency.
📝 Abstract
We study the problem of identifying the dominant arm in multi-armed bandits, where the objective is to find the action with the highest probability of exceeding the realized rewards of all other actions. Conventional mean-based and pairwise comparison-based algorithms often fail to identify the arm with the highest realized reward. To address this challenge, we introduce a novel dominant arm criterion and an efficient estimator with theoretical guarantees. Our approach relies on two key technical innovations: (i) a dominance score criterion that an arm beats the locally dominant over the partitioned reward space and (ii) a joint mixing and recycling mechanism coupled with a doubly robust estimator that guarantees simultaneous convergence of the empirical distribution functions for all arms. These key innovations pave a way to efficient computation of global arm dominance. Our proposed elimination algorithm identifies the best dominant arm with nearly optimal rate of sample complexity. Numerical experiments demonstrate that our algorithm consistently achieves exact recovery of the true dominant arm, outperforming existing baselines.