Provably Efficient Reinforcement Learning in Continuous-Time Episodic MDPs with Poisson Decision Epochs

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过扩展UCRL和Q-learning算法,解决了在具有泊松决策时段的连续时间MDP中实现高效强化学习的问题,证明了模型的最优性。
📝 Abstract
Many real-world reinforcement learning (RL) problems evolve in continuous time, where decisions occur at irregular, event-driven intervals rather than at fixed discrete steps. We study episodic continuous-time Markov Decision Processes (MDPs) in which decision epochs are governed by a homogeneous Poisson process and the reward and transition dynamics vary smoothly over time. We consider both a fixed number of jumps per episode and a fixed time budget with a random number of Poisson decision epochs. Under a Lipschitz continuity assumption in time, we exploit local smoothness through discretization and extend both UCRL (Auer and Ortner 2006) and Q-learning (Jin et al. 2018) to this setting, proving $\widetilde{O}(T^{2/3})$ regret bounds for both model-based and model-free algorithms. Finally, we establish matching $\widetildeΩ(T^{2/3})$ minimax lower bounds, showing that the rate is optimal up to logarithmic factors. These results provide the first tight regret guarantees for Lipschitz-smooth continuous-time episodic MDPs with Poisson decision epochs.
Problem

Research questions and friction points this paper is trying to address.

Continuous-time MDPs
Poisson Decision Epochs
Reinforcement Learning
Lipschitz Continuity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continuous-time MDPs
Poisson decision epochs
Lipschitz continuity
Regret bounds
UCRL and Q-learning extensions
🔎 Similar Papers
No similar papers found.
K
Kenny Guo
Department of Economics, Yale University, New Haven, Connecticut, USA
Valentio Iverson
Valentio Iverson
Undergraduate Student at University of Waterloo
Theoretical machine learningDeep Learning TheoryStatisticsOptimizationOnline Learning
S
Sahan Wijetunga
Department of Mathematics, University of California, Los Angeles, Los Angeles, California, USA
W
William Chang
Department of Mathematics, University of California, Los Angeles, Los Angeles, California, USA