Score
Designs and implements online decision systems that schedule and select actions or behaviors using contextual multi-armed bandit models, including specifying context representations, reward signals, and exploration–exploitation policies. Builds behavior-bandit schedulers that adaptively choose among candidate policies or behaviors and analyzes their empirical performance (e.g., regret or cumulative reward) under changing contexts.
This paper studies the delay-aware contextual bandit problem: a learner dynamically selects a subset of arms under stochastic contexts, where each selection incurs a random delay and a context-dependent arm request cost; the total time horizon is determined by the cumulative delays of selected subsets. The problem is formulated as a semi-Markov decision process (SMDP), the first to jointly model context dependence, arm-specific request costs, and stochastic delays. We propose an online algorithm derived from the Bellman optimality equation and establish a tight $O(sqrt{T})$ regret bound under a realizability assumption—matching the optimal rate of classical contextual bandits. The theoretical analysis is rigorous, and extensive experiments on synthetic benchmarks and a movie recommendation dataset validate both the algorithm’s empirical effectiveness and its theoretical guarantees.
This paper addresses the exploration-exploitation trade-off in contextual multi-armed bandits. We propose EE-Net, a dual-neural-network architecture: one network models the reward function for efficient exploitation, while the other directly learns instance-dependent exploration gains—bypassing conventional statistical confidence bounds—to enable adaptive exploration. To our knowledge, this is the first work to explicitly model exploration gains using neural networks. Theoretical analysis establishes an instance-dependent regret upper bound of $ ilde{O}(sqrt{T})$. Empirical evaluation on multiple real-world datasets demonstrates that EE-Net significantly outperforms both linear and state-of-the-art neural contextual bandit baselines, validating its modeling flexibility and generalization capability.
This paper studies contextual bandits with stage-wise constraints, requiring each decision to satisfy constraints simultaneously under both high-probability and expectation-based feasibility criteria, while maximizing cumulative reward and ensuring real-time constraint satisfaction. We propose, for the first time, a differentiated scaling mechanism that models confidence set radii separately for rewards and costs. A unified framework is developed to handle single or multiple constraints—whether linear or nonlinear in structure. Leveraging UCB-style exploration, eluder dimension analysis, and joint dual-constraint modeling, we establish an optimal $ ilde{O}(sqrt{T})$ regret bound and provide matching upper and lower bounds. The theoretical analysis is rigorous, and empirical simulations consistently validate the theoretical guarantees. The algorithm is scalable and applicable to complex, nonlinear constraint settings.
This paper studies the contextual combinatorial bandit problem with a dynamically evolving base arm set over time, aiming to maximize cumulative reward. To address the dual challenges of time-varying feasible action sets and context-dependent rewards, we introduce Gaussian process (GP) modeling into this framework for the first time, proposing the O’CLOK-UCB algorithm and its sparse GP-accelerated variant. Our method integrates kernelized UCB, combinatorial feasibility constraints, and Lipschitz continuity analysis. We establish a sublinear regret bound of Õ(√(λ∗(K)KTγ_T)), where λ∗(K) is the largest eigenvalue of the action covariance matrix and γ_T is the maximum information gain—revealing their coupled impact on regret. Empirical evaluation on real-world datasets demonstrates significant improvements over existing UCB-based approaches, confirming both theoretical rigor and practical efficacy.
This paper studies the nonparametric contextual bandit problem under batch constraints, where the reward function is an unknown smooth function of covariates and the policy is updated only upon completion of each batch. To address the exploration–exploitation trade-off inherent in batched learning, we propose a dynamic binning mechanism: bin widths adaptively scale with batch sizes, integrated with nonparametric regression and minimax analysis to achieve efficient estimation within the batched learning framework. We theoretically establish that only a constant number of policy updates suffice to attain the optimal online regret bound—up to logarithmic factors—and provide a matching lower bound. This is the first work to rigorously demonstrate performance equivalence between batched and online learning in the nonparametric setting, significantly reducing update frequency while enhancing practical deployability.
This work proposes a generalized multi-armed bandit and stopping problem framework grounded in behavioral preferences, departing from conventional modeling paradigms that rely on predefined states, rewards, and transition dynamics. Starting from the decision maker’s preferences over local temporal plans, the authors develop a generalized stopping representation through behavioral axioms and introduce a calendar-time cross-plan pricing mechanism. Under compact time constraints, this approach yields a rested-bandit model exhibiting index optimality. The key innovation lies in interpreting the index as the shadow price of advancing a local clock, thereby unifying—within a single preference-based framework—a diverse array of decision models, including expected utility, learning, robust, rank-dependent, Choquet, and Pandora’s box formulations. This constitutes the first theoretical foundation for index policies rooted entirely in preference theory.
This work addresses the challenge of online recommendation under heterogeneous user preferences, non-stationary context distributions, and the requirement to consistently outperform a baseline policy. The problem is formulated as a linear contextual multi-armed bandit with non-stationary heteroscedastic noise. We propose the first algorithm that simultaneously handles preference heterogeneity, context drift, and baseline constraints by extending the MED strategy to the linear setting, incorporating variance-aware suboptimality gap estimation and a constraint violation control mechanism. Theoretical analysis establishes an instance-dependent regret bound of Õ(κ/Δ̃·d²·log T) and an expected number of constraint violations bounded by Õ(d). Empirical results demonstrate that the proposed method significantly outperforms conservative baselines that ignore either context drift or preference heterogeneity.
研究解决了在线决策中动态环境和策略互动的问题,提出OnGameLearn算法,结合上下文信息进行多玩家在线游戏的学习,平衡探索与利用。
This work addresses the problem of achieving optimal decision-making in linear contextual bandits under extremely sparse parameter updates—specifically, only $O(\log\log T)$ times over horizon $T$. The paper proposes two efficient algorithms, BLCE-G and BLCE, both built upon a static scheduling mechanism that is applicable to both small and large action sets and extends naturally to generalized linear models. The key contribution lies in establishing, for the first time, a minimax-optimal regret bound under such infrequent update constraints, up to polylogarithmic factors in $T$. Notably, BLCE eliminates the need for approximate G-optimal design traditionally used in related methods, thereby substantially reducing computational complexity and emerging as the most computationally efficient algorithm to date that attains minimax optimality in this setting.
This work addresses the problem of identifying an $\varepsilon$-optimal policy in stochastic contextual bandits with $s$-sparse rewards. The authors propose an algorithm whose sample complexity depends only on the sparsity level $s$, rather than high-degree polynomials of the action space size $|A|$. By integrating information-theoretic analysis based on the Decision-Estimation Coefficient (DEC) with low-variance exploration techniques, the method applies to policy classes with bounded Natarajan dimension and extends to combinatorial semi-bandit settings. The resulting sample complexity is $\widetilde{O}\big((s/\varepsilon^2 + |A|/\varepsilon) \log(|\Pi|/\delta)\big)$, which is near-optimal up to logarithmic factors. This significantly improves upon prior results that scaled with $|A|^9$ and provides the first tight upper bound featuring no higher-order dependence on $|A|$.