Accuracy of Discretely Sampled Stochastic Policies in Continuous-time Reinforcement Learning

📅 2025-03-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In continuous-time reinforcement learning, the discrete-time execution and performance evaluation of stochastic policies have long lacked rigorous theoretical foundations. This work introduces a piecewise-constant control framework and establishes, for the first time, the weak convergence of discretely sampled policies to their continuous-time stochastic counterparts in the fine-mesh limit. We derive the optimal first-order convergence rate and provide both high-probability and almost-sure convergence guarantees. Leveraging tools from stochastic analysis and weak convergence theory, we quantify bias and variance bounds for policy evaluation and policy gradient estimation under discrete-time observations. These results furnish a rigorous theoretical basis for exploratory stochastic control. The study bridges a critical gap in the convergence analysis of policy discretization in continuous-time RL, thereby enhancing the interpretability and reliability of algorithm design.

Technology Category

Reasoning under Uncertainty: Stochastic OptimizationMachine Learning: Reinforcement LearningPlanning, Routing, and Scheduling: Mixed Discrete/Continuous Planning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsEconomics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystems
📝 Abstract
Stochastic policies are widely used in continuous-time reinforcement learning algorithms. However, executing a stochastic policy and evaluating its performance in a continuous-time environment remain open challenges. This work introduces and rigorously analyzes a policy execution framework that samples actions from a stochastic policy at discrete time points and implements them as piecewise constant controls. We prove that as the sampling mesh size tends to zero, the controlled state process converges weakly to the dynamics with coefficients aggregated according to the stochastic policy. We explicitly quantify the convergence rate based on the regularity of the coefficients and establish an optimal first-order convergence rate for sufficiently regular coefficients. Additionally, we show that the same convergence rates hold with high probability concerning the sampling noise, and further establish a $1/2$-order almost sure convergence when the volatility is not controlled. Building on these results, we analyze the bias and variance of various policy evaluation and policy gradient estimators based on discrete-time observations. Our results provide theoretical justification for the exploratory stochastic control framework in [H. Wang, T. Zariphopoulou, and X.Y. Zhou, J. Mach. Learn. Res., 21 (2020), pp. 1-34].
Problem

Research questions and friction points this paper is trying to address.

Challenges in executing stochastic policies in continuous-time reinforcement learning.
Convergence of controlled state process with discrete policy sampling.
Analysis of bias and variance in policy evaluation estimators.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Discrete sampling of stochastic policies in continuous-time RL.
Piecewise constant controls for policy execution framework.
Quantified convergence rates for controlled state processes.
🔎 Similar Papers
No similar papers found.
Y
Yanwei Jia
Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong, Hong Kong
O
Ouyang Du
Department of Mathematical Sciences, Tsinghua University, China
Y
Yufei Zhang
Department of Mathematics, Imperial College London, United Kingdom