Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivit

๐Ÿ“… 2026-09-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the K-armed bandit problem with A optimal arms, where existing literature exhibits theoretical gaps in minimax regret optimization and non-adaptivity under multiple optimal arms. By refining the analytical framework of subsampling algorithms and integrating minimax theory with asymptotic complexity derivations, this work provides a rigorous characterization of regret bounds across the full range 1 โ‰ค A โ‰ค K โˆ’ 1. It establishes a tighter regret upper bound of O(((Kโˆ’A)/โˆš(KA))ยทโˆšT) with a matching lower bound, yielding a near minimax-optimal regret rate. Furthermore, it proves that prior knowledge of the number of optimal arms A is theoretically necessary for adaptation. These contributions deliver a complete characterization of the multiple optimal arms setting, substantially improving upon existing results in the literature.
๐Ÿ“ Abstract
We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers. For $K$-armed bandits with $A$ optimal arms, we first provide a sharper analysis of previous sub-sampling algorithms (De Heide et al., 2021; Zhu and Nowak, 2020), establishing a $\tilde{O}\Big(\frac{K-A}{\sqrt{KA}}\sqrt{T} \Big)$ minimax regret, where $T$ is the total number of interactions and $\tilde O(\cdot)$ drops all constant and logarithmic factors, improving the previous $\tilde{O}(\sqrt{KT/A})$ regret. We then provide a matching lower bound up to logarithmic factors, indicating that our established rate is nearly minimax-optimal. We further show that the knowledge of $A$ up to $\tilde{O}(1)$ factors is necessary to achieve near-optimal regret, as near-optimal algorithms for one number of optimal arms must incur substantially larger regret than optimal regret for a smaller number. Overall, our results provide a comprehensive minimax characterization of $K$-armed bandits with $A$ over the entire range of $1 \leq A \leq K-1$.
Problem

Research questions and friction points this paper is trying to address.

Multi-armed bandits
Multiple optimal arms
Minimax regret
Lower bound
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-armed bandits
Minimax regret
Multiple optimal arms
Sub-sampling algorithms
Non-adaptivity
๐Ÿ”Ž Similar Papers
No similar papers found.