🤖 AI Summary
This work addresses the challenge in finite-time multi-armed bandits where existing algorithms struggle to simultaneously satisfy practical constraints and achieve strong theoretical performance. The paper proposes a class of regularized greedy algorithms and, for the first time, derives a decomposable form of their finite-time regret bounds, thereby revealing the trade-off between exploration cost and convergence error. A calibration criterion for the regularization parameter is established based on this analysis. Under a Bernoulli reward model and using exponential decay convergence arguments, the proposed algorithm consistently matches or outperforms state-of-the-art methods across extensive experiments, while also providing tighter regret guarantees for classical greedy strategies.
📝 Abstract
Organizations increasingly rely on sequential experimentation to improve decision-making. While the multi-armed bandit literature has developed algorithms with strong asymptotic regret guarantees, many practical applications operate over finite and externally imposed horizons. Motivated by the finite-horizon setting, we develop a class of regularized greedy algorithms for multi-armed Bernoulli bandits. We derive the first finite-horizon regret envelopes for regularized greedy bandits, showing that finite-horizon regret decomposes into transient exploration costs and a suboptimal convergence term that decays exponentially with the regularization strength. This characterization yields principled calibration rules for the regularization parameters and, as a limiting case, sharper regret guarantees for the classical greedy policy. Across extensive numerical experiments, calibrated regularized greedy policies consistently match or outperform state-of-the-art algorithms. These results suggest that regularized greedy policies can provide an effective approach for finite-horizon bandit problems.