Variance-Aware Fine-Grained Gap-Dependent Bounds for Online Reinforcement Learning

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of fine-grained gap-dependent bounds and the excessive policy switching costs in variance-aware online reinforcement learning. Building upon the UCB-Bernstein+ algorithm, this work introduces a refined variance-based exploration bonus for the first time, bridging a theoretical gap in fine-grained analysis, and proposes a phased policy update framework. The research establishes the first fine-grained gap-dependent regret upper bound, achieves optimal local policy switching costs, and significantly improves worst-case performance guarantees. Comprehensive experiments thoroughly validate the effectiveness and superiority of the proposed approach.
📝 Abstract
We study model-free online reinforcement learning (RL) for episodic tabular Markov decision processes, focusing on both gap-dependent regret and policy switching cost. While fine-grained gap-dependent analysis has been established for model-free RL algorithms using Hoeffding-type exploration bonuses, such results for model-free algorithms with variance-based exploration bonuses remain unknown, despite their superior worst-case and coarse-grained gap-dependent guarantees. In this paper, we resolve this open problem by establishing the first fine-grained gap-dependent regret upper bound for UCB-Bernstein+, a refined UCB-Bernstein algorithm, in variance-aware model-free online RL. Moreover, by integrating a stage-wise policy update design into our fine-grained framework and using refined variance-based bonuses, we achieve the best-known gap-dependent local switching cost to date. In addition, our analysis yields improved worst-case guarantees for both regret and local switching cost over the original UCB-Bernstein algorithm. Numerical experiments further demonstrate that UCB-Bernstein+ achieves favorable empirical performance in both regret and local switching cost.
Problem

Research questions and friction points this paper is trying to address.

online reinforcement learning
gap-dependent regret
variance-aware exploration
policy switching cost
model-free RL
Innovation

Methods, ideas, or system contributions that make the work stand out.

Variance-Aware
Fine-Grained Gap-Dependent Bounds
Model-Free Online Reinforcement Learning
Policy Switching Cost
UCB-Bernstein
🔎 Similar Papers
No similar papers found.