Linear Bandits under Exact Sliding-Window Constraints

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of optimal decision-making and online learning in linear bandits under strict sliding window constraints. It introduces the transition diameter to quantify state reachability and constructs a finite-memory control model. A novel algorithmic framework is proposed that integrates optimistic remaining-horizon planning with infrequent policy switching, enabling low-frequency policy updates based on state history. Extensive evaluations on both real-world and synthetic benchmarks demonstrate that the proposed algorithm strictly satisfies sliding window feasibility constraints while achieving cumulative rewards and regret bounds comparable to existing baselines, with significantly fewer policy updates. These results confirm that the approach effectively balances rigorous constraint satisfaction with efficient online learning.
📝 Abstract
We study linear bandits under exact sliding-window constraints, where every consecutive block of actions must belong to a prescribed feasible set. In the offline setting, where the reward function is known, we show that convexity and cyclic-shift invariance make a stationary solution optimal when $w\mid T$ and within an additive $O(w)$ gap otherwise. In the online setting, we show that geometric structure alone is insufficient for learning, and sublinear regret can be impossible. We introduce a transition diameter $τ$ that quantifies feasible reachability and develop a rare-switching OFUL algorithm with regret $\widetilde{O}(d\sqrt{T}+τd+w)$ against the offline-optimal feasible trajectory. Finally, we remove cyclic invariance and consider general sliding-window constraints, where optimal behavior may be non-stationary. We represent recent action history as the state of a finite-memory control problem and introduce a history-state diameter $D$ that measures feasible communication between viable histories. Combining optimistic remaining-horizon planning with rare policy updates, we obtain a regret bound of $\widetilde{O}(d\sqrt{T}+dD+w)$. We evaluate our approach on real-world and synthetic benchmarks, showing that it maintains exact feasibility while achieving reward and regret comparable to baselines with substantially fewer policy updates.
Problem

Research questions and friction points this paper is trying to address.

Linear Bandits
Sliding-Window Constraints
Online Learning
Regret Minimization
Non-stationary
Innovation

Methods, ideas, or system contributions that make the work stand out.

Linear Bandits
Sliding-Window Constraints
Rare-Switching OFUL
Transition Diameter
History-State Diameter
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Seyed Mohammad Hadi Hosseini
Y
Yasin Abbasi-Yadkori
Sattar Vakili
Sattar Vakili
MediaTek Research
Machine Learning