Block Optimism for Nonstationary Bandits with Latent Linear Dynamics

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses non-stationary bandit problems with latent linear dynamics, where action-induced perturbations on future states complicate long-horizon planning and cause existing methods to suffer from excessively high regret bounds. To overcome this, the proposed method employs a recurrence approximation to truncate the infinite-memory reward process, constructing a finite-memory block-level surrogate model. By integrating an Upper Confidence Bound (UCB) algorithm to maintain confidence sets, it achieves adaptive block-level optimistic exploration for optimizing decision sequences. This work establishes the first $\tilde{O}(\sqrt{T})$ regret bound for bilinear reward observations under an open-loop baseline, breaking through the performance bottleneck of conventional uniform exploration strategies and reducing cumulative regret from $\tilde{O}(T^{2/3})$ to the optimal theoretical limit.
📝 Abstract
We study an endogenous nonstationary stochastic bandit problem with latent linear dynamics, where actions affect both immediate rewards and the future evolution of an unobserved latent state. Rewards are bilinear in the current action and latent state, inducing history-dependent rewards and a nontrivial long-horizon planning problem. The existing explore-then-commit approach achieves $\tilde{O}(T^{2/3})$ regret by uniformly exploring to estimate the latent dynamics and then committing to an optimized open-loop action sequence. We show that this rate can be improved via adaptive block-level optimism. Our key step is a cyclic approximation: under stable dynamics, the infinite-memory reward process can be truncated, and the open-loop benchmark can be approximated by optimizing a finite-memory block-level proxy. Building on this reduction, we propose a UCB-based block algorithm that maintains confidence sets for the truncated dynamics parameters and selects blocks optimistically. We prove a regret bound of order $\tilde{O}(\sqrt T)$, significantly improving over the previous $\tilde{O}(T^{2/3})$ guarantee for the same model. To the best of our knowledge, this is the first $\tilde{O}(\sqrt T)$ regret guarantee for latent linear-dynamics bandits with bilinear reward observations and an open-loop action-sequence benchmark.
Problem

Research questions and friction points this paper is trying to address.

nonstationary bandits
latent linear dynamics
bilinear rewards
regret minimization
long-horizon planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

nonstationary bandits
latent linear dynamics
block-level optimism
regret bound
bilinear rewards
🔎 Similar Papers
2024-02-05arXiv.orgCitations: 1
💼 Related Jobs
No related jobs found.
T
Taehyun Hwang
Seoul National University
H
Hyunjun Choi
Seoul National University
H
Heesang Ann
Seoul National University
Min-hwan Oh
Min-hwan Oh
Seoul National University
Reinforcement LearningBandit AlgorithmsMachine Learning