🤖 AI Summary
This paper studies threshold-type multi-armed bandit (MAB) problems with mean monotonicity, aiming to identify the first arm, the *k*-th arm, or the arm whose mean reward is closest to a given threshold τ. To address this novel paradigm, we first derive a threshold-dependent asymptotic regret lower bound, revealing that only arms in the neighborhood of τ govern the fundamental performance limit—thereby extending classical MAB theory. Second, we design an efficient online algorithm that explicitly exploits both monotonicity and threshold constraints, and prove it achieves the derived lower bound asymptotically. Finally, Monte Carlo simulations validate the algorithm’s optimality and robustness in practical applications, including CQI allocation in wireless communications and clinical dose-finding trials. This work establishes a new structured MAB model, provides the first threshold-aware information-theoretic regret bound, and delivers a provably optimal algorithm—advancing sequential decision-making under structural constraints.
📝 Abstract
In multi-armed bandit problems, the typical goal is to identify the arm with the highest reward. This paper explores a threshold-based bandit problem, aiming to select an arm based on its relation to a prescribed threshold (τ). We study variants where the optimal arm is the first above (τ), the (k^{th}) arm above or below it, or the closest to it, under a monotonic structure of arm means. We derive asymptotic regret lower bounds, showing dependence only on arms adjacent to (τ). Motivated by applications in communication networks (CQI allocation), clinical dosing, energy management, recommendation systems, and more. We propose algorithms with optimality validated through Monte Carlo simulations. Our work extends classical bandit theory with threshold constraints for efficient decision-making.