Learning When to Update: A Near-Optimal Timing Bandit Approach

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of determining optimal update timings for resource-intensive systems, such as machine learning models, operating under unknown degradation patterns in dynamic environments, where a fundamental trade-off exists between update costs and performance loss. To this end, this work formulates the update timing problem as a novel temporal bandit paradigm. By exploiting structural properties of intermediate degradation revealed over extended intervals, it proposes the BCAE and OE-BCAE algorithms, which integrate successive arm elimination with optimism-enhanced confidence bounds. Theoretically, the proposed algorithms achieve an $\tilde{O}(\sqrt{T})$ regret bound, surpassing the established lower bound of standard bandits in $K$-arm settings. Empirical simulations further demonstrate that these methods maintain low regret and high stability across varying numbers of arms, significantly outperforming existing approaches.
πŸ“ Abstract
Systems operating in dynamic environments require timely updates to sustain performance. For resource-intensive systems such as machine learning models and digital twins, strategically timing updates is essential. Updating too frequently wastes resources, while updating too infrequently leads to costly performance degradation. The problem is particularly challenging when the system's degradation pattern is unknown a priori, as is common in new operating environments. We formalize this challenge as a novel \emph{timing bandit} problem, where each arm represents a candidate update interval with a fixed update cost and an unknown, stochastic degradation cost. Three structural properties distinguish this setting from standard multi-armed bandits: selecting an interval commits the learner to multiple time slots before the next update; arm costs are composed of per-step degradation costs and a fixed update cost; and selecting a longer interval naturally reveals degradation at every intermediate step, providing consecutive feedback relevant to shorter intervals. By exploiting these structures, we develop Balanced Consecutive Arm Elimination (BCAE). BCAE achieves $\tilde{O}(\sqrt{T})$ regret, improving upon the $\tildeΞ©(K\sqrt{T})$ regret of standard bandit algorithms in this setting, where $K$ is the number of candidate update intervals. We further propose an Optimism-Enhanced variant (OE-BCAE) that integrates lower-confidence-bound principles to improve empirical adaptivity while preserving the same regret order. Moreover, the regret bound achieved by our algorithms matches the theoretical lower bound up to logarithmic factors. Simulation results demonstrate that our algorithms achieve low regret and remain stable as both the number of arms and the update cost vary.
Problem

Research questions and friction points this paper is trying to address.

timing bandit
update scheduling
dynamic environments
performance degradation
multi-armed bandits
Innovation

Methods, ideas, or system contributions that make the work stand out.

Timing Bandit
Balanced Consecutive Arm Elimination
Regret Bound
Update Scheduling
Optimism-Enhanced
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Q
Qiulin Lin
City University of Hong Kong, Hong Kong, China
J
Junyan Su
City University of Hong Kong, Hong Kong, China
Liyuan Wang
Liyuan Wang
Tsinghua University
bio-inspired learningcontinual learningneuroscience
M
Minghua Chen
The Chinese University of Hong Kong (Shenzhen), China; City University of Hong Kong, Hong Kong, China