Clus-UCB: A Near-Optimal Algorithm for Clustered Bandits

📅 2025-08-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper studies the stochastic multi-armed bandit problem with known cluster structure: arms are partitioned into clusters such that the expected rewards of arms within each cluster differ by at most a given threshold. This setting models real-world applications—such as online advertising and clinical trials—where rewards are influenced by multiple factors and exhibit natural groupings. Since the classical Lai–Robbins lower bound is not tight under this structure, we propose Clus-UCB, the first algorithm integrating the KL-UCB framework with intra-cluster reward dependence modeling. Clus-UCB introduces a novel cluster-level confidence-bound index mechanism that enables information sharing among arms in the same cluster. We establish a strictly tighter asymptotic regret lower bound for Clus-UCB than the classical one. Empirical evaluations demonstrate that Clus-UCB significantly outperforms KL-UCB and other baselines across diverse structured dependency scenarios.

Technology Category

Machine Learning: Online Learning & BanditsReasoning under Uncertainty: Stochastic OptimizationIntelligent Robots: Learning & Optimization for ROB

Application Category

Economics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystemsGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsUser Modeling, Personalization and Recommendation: User modeling for targeted and personalized online advertising
📝 Abstract
We study a stochastic multi-armed bandit setting where arms are partitioned into known clusters, such that the mean rewards of arms within a cluster differ by at most a known threshold. While the clustering structure is known a priori, the arm means are unknown. This framework models scenarios where outcomes depend on multiple factors -- some with significant and others with minor influence -- such as online advertising, clinical trials, and wireless communication. We derive asymptotic lower bounds on the regret that improve upon the classical bound of Lai & Robbins (1985). We then propose Clus-UCB, an efficient algorithm that closely matches this lower bound asymptotically. Clus-UCB is designed to exploit the clustering structure and introduces a new index to evaluate an arm, which depends on other arms within the cluster. In this way, arms share information among each other. We present simulation results of our algorithm and compare its performance against KL-UCB and other well-known algorithms for bandits with dependent arms. Finally, we address some limitations of this work and conclude by mentioning possible future research.
Problem

Research questions and friction points this paper is trying to address.

Optimizing clustered bandits with known cluster structure
Improving regret bounds in multi-armed bandit problems
Exploiting arm dependencies within clusters for efficient learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Clustered bandit algorithm with known clusters
New index for arms within clusters
Asymptotic near-optimal regret performance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Aakash Gore
Department of Electrical Engineering, Indian Institute Of Technology Bombay
Prasanna Chaporkar
Prasanna Chaporkar
Electrical Engineering, IIT Bombay