delayed-switching bandits

Designs and analyzes sequential decision-making models and algorithms for multi-armed bandit problems in which switching between actions incurs costs and actions or feedback are delayed, including settings with multiple simultaneous plays. Work includes explicitly modeling stochastic delays, building policies (e.g., UCB-style replacement/selection rules) that trade off switching costs and delay, and proving regret bounds (upper bounds and matching lower bounds).

delayed-switchingbandits

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Biased Dueling Bandits with Stochastic Delayed Feedback

Aug 26, 2024
BY
Bongsoo Yi
🏛️ University of North Carolina at Chapel Hill | University of California, Davis

This work studies the preference-biased dueling bandits problem with stochastic delayed feedback, addressing the core challenge of delayed reward feedback hindering timely policy updates in online recommendation and advertising. We are the first to model the coupled effect of delay and preference bias, and propose two adaptive algorithms: one for settings with known delay distribution, and another for more realistic scenarios where only the expected delay is available. Methodologically, we integrate a delay-aware UCB framework, bias-corrected pairwise comparison estimation, and rigorous regret analysis. Theoretically, both algorithms achieve the optimal $O(sqrt{T})$ regret bound—matching that of the non-delayed setting—and constitute the first delay-robust optimal algorithms for this problem. Empirical evaluation on synthetic and real-world datasets validates the tightness of our theoretical bounds and demonstrates significant performance gains over baselines.

Addressing biased dueling bandits with delayed feedbackDeveloping algorithms for unknown delay distributionsHandling stochastic delays in action feedback

Stochastic Submodular Bandits with Delayed Composite Anonymous Bandit Feedback

Mar 23, 2023
MP
Mohammad Pedramfar
🏛️ McGill University | Mila - Quebec Artificial Intelligence Institute | Purdue University

This paper studies the combinatorial multi-armed bandit (CMAB) problem with stochastic submodular expected rewards and delayed composite anonymous feedback—addressing practical challenges including feedback aliasing, indistinguishable sources, and delayed arrival. We propose an online learning algorithm based on greedy sampling and decoupled estimation of delayed feedback. For the first time, we establish a unified regret bound of $ ilde{O}(T^{2/3} + T^{1/3} u)$ under three general delay models: bounded adversarial, stochastically independent, and stochastically conditionally independent delays—revealing a universal additive impact of delay on performance. Theoretically, this bound strictly improves upon existing full-feedback delay methods. Empirically, our algorithm significantly reduces cumulative regret on both synthetic benchmarks and real-world submodular tasks, including influence maximization.

Anonymized RewardsCombinatorial Multi-Armed BanditDelayed Feedback

Lipschitz Bandits with Stochastic Delayed Feedback

Sep 30, 2025
ZL
Zhongxuan Liu
🏛️ University of California, Davis

This paper studies the Lipschitz bandit problem under stochastic delayed feedback: the action space is a metric space with Lipschitz-continuous expected rewards, and reward observations incur i.i.d. random delays—either bounded or unbounded. To address this novel setting, we propose a delay-aware zooming algorithm for bounded delays and a phased learning strategy for unbounded delays, achieving the first sublinear regret bounds of $O(T^{(d+1)/(d+2)}log T)$ and $O(T^{(d+2)/(d+3)}log T)$, respectively, where $d$ denotes the metric space dimension. Our theoretical analysis establishes matching lower bounds, confirming near-optimality. Experiments demonstrate robustness and efficiency across diverse delay distributions. The core innovation lies in tightly coupling Lipschitz structure with dynamic modeling of delay effects, enabling principled continuous decision-making under information lag.

Achieving sublinear regret guarantees under delayed observationsAlgorithms for bounded and unbounded stochastic delay settingsLipschitz bandits with stochastic delayed reward feedback

This paper addresses the problem of safely and efficiently improving online learning performance for stochastic multi-armed bandits (MAB) and combinatorial MAB under offline–online distribution shift. To tackle distribution mismatch between offline and online data, we propose MIN-UCB—a novel adaptive algorithm that automatically discards low-quality offline data when the shift magnitude is unknown, ensuring safety, and actively leverages bounded-shift offline information to substantially reduce regret when the shift is bounded. We establish, for the first time, a theoretical lower bound on regret for MAB with biased offline data. We prove that MIN-UCB achieves tight instance-dependent and instance-independent regret bounds—strictly outperforming classical UCB. Both theoretical analysis and numerical experiments demonstrate its robustness and superiority across diverse shift regimes.

Addressing distribution mismatch between offline data and online rewards.Developing adaptive policies that selectively use informative offline data.Leveraging offline data to enhance online learning in bandit problems.

This work addresses the challenge of unbounded metric complexity in online convex optimization with noisy feedback, where the objective involves jointly minimizing high-dimensional dynamic quadratic hitting costs and ℓ₂-norm switching costs. To tackle this problem under stochastic environments with unknown hitting cost structures, the authors propose the SCaLE algorithm, which achieves, for the first time, a sublinear dynamic regret guarantee for such joint settings. By introducing spectral regret analysis, SCaLE effectively disentangles regret contributions arising from eigenvalue estimation errors and perturbations in the eigenvector basis. Theoretical analysis establishes a distribution-free sublinear dynamic regret bound for SCaLE, and empirical evaluations demonstrate its superior performance and statistical consistency against multiple baselines.

bandit feedbackdynamic regretmetric movement cost

Latest Papers

What's happening recently
View more

This study addresses the problem of optimal decision-making in stochastic linear bandits under delayed feedback, systematically distinguishing for the first time among three delay models: loss-independent, loss-dependent, and “delay-as-reward.” Through regret analysis and matching upper and lower bound constructions, it elucidates the interplay between linear structure and delayed feedback: loss-independent delays incur only a dimension-independent additive penalty, whereas loss-dependent delays induce a penalty proportional to the square root of the feature dimension. The work further demonstrates that several classical results from multi-armed bandits fail to extend to the linear setting. The derived regret bounds are nearly optimal across multiple delay models, significantly improving upon existing results—particularly in the loss-independent case.

delayed feedbackloss-dependent delaysmulti-armed bandits

This work proposes a generalized multi-armed bandit and stopping problem framework grounded in behavioral preferences, departing from conventional modeling paradigms that rely on predefined states, rewards, and transition dynamics. Starting from the decision maker’s preferences over local temporal plans, the authors develop a generalized stopping representation through behavioral axioms and introduce a calendar-time cross-plan pricing mechanism. Under compact time constraints, this approach yields a rested-bandit model exhibiting index optimality. The key innovation lies in interpreting the index as the shadow price of advancing a local clock, thereby unifying—within a single preference-based framework—a diverse array of decision models, including expected utility, learning, robust, rank-dependent, Choquet, and Pandora’s box formulations. This constitutes the first theoretical foundation for index policies rooted entirely in preference theory.

Behavioral PreferencesDecision TheoryMulti-Armed Bandits

This work addresses the challenge of minimizing regret against arbitrarily switching mixed strategies of forgetful opponents in extensive-form games under bandit feedback. To this end, we propose the first online learning algorithm that simultaneously achieves low switching regret and high computational efficiency. By leveraging the game tree structure and incorporating an adaptive parameter mechanism, our method effectively controls the frequency of strategy switches while optimizing the regret bound. The algorithm attains a switching regret bound of Õ((1/ρ + ρK)√(HAT)) and operates with a per-round time complexity of only O(HB), significantly enhancing both scalability and practical applicability.

bandit problemextensive-form gameoblivious adversary

Hot Scholars

DS

David Stap

NXAI
Machine TranslationMachine LearningNatural Language Processing
SB

Sebastian Böck

NXAI
Deep LearningMusic Information RetrievalAudio Signal ProcessingComputer Vision
SH

Sepp Hochreiter

Institute for Machine Learning, Johannes Kepler University Linz
Machine LearningDeep LearningArtificial IntelligenceNeural Networks
GK

Günter Klambauer

Prof., LIT AI Lab & Institute for Machine Learning, Johannes Kepler University
machine learningchemoinformaticsdeep learningdrug design
TS

Thomas Schmied

PhD Student, Institute for Machine Learning, Johannes Kepler University Linz
Deep LearningReinforcement LearningNatural Language ProcessingContinual Learning