🤖 AI Summary
This study addresses the challenges of optimal decision-making and online learning in linear bandits under strict sliding window constraints. It introduces the transition diameter to quantify state reachability and constructs a finite-memory control model. A novel algorithmic framework is proposed that integrates optimistic remaining-horizon planning with infrequent policy switching, enabling low-frequency policy updates based on state history. Extensive evaluations on both real-world and synthetic benchmarks demonstrate that the proposed algorithm strictly satisfies sliding window feasibility constraints while achieving cumulative rewards and regret bounds comparable to existing baselines, with significantly fewer policy updates. These results confirm that the approach effectively balances rigorous constraint satisfaction with efficient online learning.
📝 Abstract
We study linear bandits under exact sliding-window constraints, where every consecutive block of actions must belong to a prescribed feasible set. In the offline setting, where the reward function is known, we show that convexity and cyclic-shift invariance make a stationary solution optimal when $w\mid T$ and within an additive $O(w)$ gap otherwise. In the online setting, we show that geometric structure alone is insufficient for learning, and sublinear regret can be impossible. We introduce a transition diameter $τ$ that quantifies feasible reachability and develop a rare-switching OFUL algorithm with regret $\widetilde{O}(d\sqrt{T}+τd+w)$ against the offline-optimal feasible trajectory. Finally, we remove cyclic invariance and consider general sliding-window constraints, where optimal behavior may be non-stationary. We represent recent action history as the state of a finite-memory control problem and introduce a history-state diameter $D$ that measures feasible communication between viable histories. Combining optimistic remaining-horizon planning with rare policy updates, we obtain a regret bound of $\widetilde{O}(d\sqrt{T}+dD+w)$. We evaluate our approach on real-world and synthetic benchmarks, showing that it maintains exact feasibility while achieving reward and regret comparable to baselines with substantially fewer policy updates.