A UCB Bandit Algorithm for General ML-Based Estimators

๐Ÿ“… 2026-01-03
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of principled exploration in sequential decision-making with complex machine learning models, which has been hindered by the absence of general concentration inequalities. The authors propose ML-UCB, the first algorithm that derives a universal concentration inequality applicable to arbitrary models by leveraging empirically characterizable learning curvesโ€”specifically, those exhibiting power-law decay in mean squared error. This approach guarantees sublinear regret without requiring model-specific theoretical analysis. ML-UCB integrates upper confidence bound (UCB) strategies with learning curve modeling, concentration bounds under power-law assumptions, and online matrix factorization to effectively balance exploration and exploitation. Experiments on a simulated two-tower collaborative filtering recommendation system demonstrate that ML-UCB significantly outperforms LinUCB, confirming its generality and efficacy.

Technology Category

Machine Learning: Online Learning & BanditsReasoning under Uncertainty: Sequential Decision MakingSearch and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: ML for personalized search and recommendationsGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
๐Ÿ“ Abstract
We present ML-UCB, a generalized upper confidence bound algorithm that integrates arbitrary machine learning models into multi-armed bandit frameworks. A fundamental challenge in deploying sophisticated ML models for sequential decision-making is the lack of tractable concentration inequalities required for principled exploration. We overcome this limitation by directly modeling the learning curve behavior of the underlying estimator. Specifically, assuming the Mean Squared Error decreases as a power law in the number of training samples, we derive a generalized concentration inequality and prove that ML-UCB achieves sublinear regret. This framework enables the principled integration of any ML model whose learning curve can be empirically characterized, eliminating the need for model-specific theoretical analysis. We validate our approach through experiments on a collaborative filtering recommendation system using online matrix factorization with synthetic data designed to simulate a simplified two-tower model, demonstrating substantial improvements over LinUCB
Problem

Research questions and friction points this paper is trying to address.

multi-armed bandit
machine learning
concentration inequalities
sequential decision-making
exploration
Innovation

Methods, ideas, or system contributions that make the work stand out.

ML-UCB
learning curve modeling
generalized concentration inequality
sublinear regret
multi-armed bandits