Expected Sample Complexity in Multi-Armed Bandits

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the incomplete characterization and analysis of expected sample complexity in stochastic multi-armed bandits. It proposes the ACE theoretical framework, which builds upon explore-then-commit ε-greedy and Thompson sampling algorithms to distinguish between known and unknown suboptimality gap settings. The framework reveals the limitations of deterministic algorithms and establishes tight lower bounds, while demonstrating that expected sample complexity implies almost sure convergence and can be translated into regret bounds. By achieving near-matching alignment between algorithmic performance and theoretical lower bounds, this work rigorously proves a performance separation between the two settings. Ultimately, it provides a unified paradigm for sample complexity analysis in stochastic bandit problems.
📝 Abstract
Sample complexity is a widely used metric in sequential decision-making problems, defined as the number of suboptimal decisions during the interaction between the agent and an environment. We study the sample complexity of stochastic multi-armed bandit problems and introduce the expected sample complexity performance measure, analyzing it in a novel framework called approximately correct in expectation (ACE). We show that ACE guarantees imply almost sure convergence to the optimal expected reward, in contrast to high-probability guarantees found in other frameworks, and also show how to convert ACE guarantees into explicit expected regret bounds. We further show that, in contrast to existing measures, deterministic algorithms cannot obtain favorable ACE bounds, and analyze stochastic algorithms in two settings: when the allowed suboptimality level $ε$ is known to the algorithm and when it is unknown. In the former, we devise an explore-then-$ε$-greedy algorithm, and in the latter, we analyze the expected sample complexity of Thompson sampling. Finally, we establish nearly matching lower bounds for both settings, showing that the algorithms are tight in $ε$ and proving a performance separation between the two regimes.
Problem

Research questions and friction points this paper is trying to address.

Multi-Armed Bandits
Sample Complexity
Expected Sample Complexity
Approximately Correct in Expectation
Sequential Decision-Making
Innovation

Methods, ideas, or system contributions that make the work stand out.

Expected Sample Complexity
Approximately Correct in Expectation (ACE)
Multi-Armed Bandits
Explore-then-epsilon-greedy
Thompson Sampling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.