🤖 AI Summary
This study addresses the challenge of boundary handling in reflected dynamic reinforcement learning within hypercube state spaces. It proposes a continuous-time deterministic policy gradient framework grounded in controlled reflected stochastic differential equations, enforcing Neumann boundary conditions through soft penalties or hard constraints. Theoretically, the work establishes a connection between the value function and the Neumann Bellman equation, introduces an advantage rate function alongside a martingale representation theorem, and develops a continuous-time deep deterministic policy gradient algorithm. Empirically, the proposed approach significantly reduces boundary residuals and enhances learning stability, with its effectiveness validated on a reservoir control task.
📝 Abstract
We develop a continuous-time deterministic policy gradient framework for reinforcement learning with reflected state dynamics, where the state process is governed by a controlled reflected stochastic differential equation on a hypercube. Under suitable regularity assumptions, we establish the connection between the value function and the Neumann Bellman equation, introduce an advantage-rate function that yields a deterministic policy gradient formula, and prove the martingale characterization theorem. Motivated by these theoretical results, we propose a continuous-time deep deterministic policy gradient algorithm for reflected stochastic systems, in which the Neumann boundary condition is imposed via either soft penalization or hard architectural constraint. We further quantify the discrepancy between the ideal continuous-time dynamics and the discretely sampled exploratory dynamics executed in practice, showing that the error decays as the time grid is refined and exploration noise vanishes. Our experiments on reservoir control problems illustrate the effectiveness of the RL framework, highlighting that boundary-aware methods substantially reduce Neumann boundary residuals and enhance learning stability.