Threshold-Based Optimal Arm Selection in Monotonic Bandits: Regret Lower Bounds and Algorithms

📅 2025-09-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper studies threshold-type multi-armed bandit (MAB) problems with mean monotonicity, aiming to identify the first arm, the *k*-th arm, or the arm whose mean reward is closest to a given threshold τ. To address this novel paradigm, we first derive a threshold-dependent asymptotic regret lower bound, revealing that only arms in the neighborhood of τ govern the fundamental performance limit—thereby extending classical MAB theory. Second, we design an efficient online algorithm that explicitly exploits both monotonicity and threshold constraints, and prove it achieves the derived lower bound asymptotically. Finally, Monte Carlo simulations validate the algorithm’s optimality and robustness in practical applications, including CQI allocation in wireless communications and clinical dose-finding trials. This work establishes a new structured MAB model, provides the first threshold-aware information-theoretic regret bound, and delivers a provably optimal algorithm—advancing sequential decision-making under structural constraints.

Technology Category

Machine Learning: Online Learning & BanditsReasoning under Uncertainty: Sequential Decision MakingMultiagent Systems: Mechanism Design

Application Category

Economics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystemsGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsResponsible Web: Human-perceived consequences of algorithmic deployment on the web
📝 Abstract
In multi-armed bandit problems, the typical goal is to identify the arm with the highest reward. This paper explores a threshold-based bandit problem, aiming to select an arm based on its relation to a prescribed threshold (τ). We study variants where the optimal arm is the first above (τ), the (k^{th}) arm above or below it, or the closest to it, under a monotonic structure of arm means. We derive asymptotic regret lower bounds, showing dependence only on arms adjacent to (τ). Motivated by applications in communication networks (CQI allocation), clinical dosing, energy management, recommendation systems, and more. We propose algorithms with optimality validated through Monte Carlo simulations. Our work extends classical bandit theory with threshold constraints for efficient decision-making.
Problem

Research questions and friction points this paper is trying to address.

Selecting optimal arms relative to a threshold τ
Deriving regret lower bounds for threshold-based bandits
Designing algorithms for monotonic bandit problems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Threshold-based optimal arm selection
Asymptotic regret lower bounds derivation
Algorithms validated via Monte Carlo simulations
🔎 Similar Papers
Indian Institute of Technology
C
Chanakya Varude
Department of Electrical Engineering, Indian Institute of Technology, Bombay
J
Jay Chaudhary
Department of Electrical Engineering, Indian Institute of Technology, Bombay
S
Siddharth Kaushik
Department of Electrical Engineering, Indian Institute of Technology, Bombay
Prasanna Chaporkar
Prasanna Chaporkar
Electrical Engineering, IIT Bombay