🤖 AI Summary
This study addresses the continuous-time optimal asset allocation problem under stochastic volatility and portfolio constraints by introducing an entropy-regularized reinforcement learning framework. Modeling the control policy as a probability distribution rather than a deterministic function, the authors derive the associated entropy-regularized Hamilton-Jacobi-Bellman (HJB) equation via dynamic programming and propose an optimal exploratory policy of truncated Gaussian form. Leveraging stochastic control theory and martingale methods, they establish the existence of solutions to the resulting nonlinear quasilinear parabolic PDE and obtain semi-closed-form expressions for both the value function and the optimal policy. Furthermore, they develop an implementable continuous-time Actor-Critic algorithm and prove the convergence of its policy improvement process, thereby revealing an intrinsic connection between entropy-regularized relaxed controls and continuous-time reinforcement learning.
📝 Abstract
We study the problem of optimal portfolio selection under stochastic volatility within a continuous time reinforcement learning framework with portfolio constraints. Exploration is modeled through entropy-regularized relaxed controls, where the investor selects probability distributions over admissible portfolio allocations rather than deterministic strategies. Using dynamic programming arguments, we derive the associated entropy-regularized Hamilton-Jacobi-Bellman equation, whose Hamiltonian involves optimization over probability measures supported on a compact control set. We show that the optimal exploratory policy takes the form of a truncated Gaussian distribution characterized by spatial derivatives of the solution of the resulting nonlinear quasilinear parabolic partial differential equation. Under suitable structural conditions on the model coefficients, we prove the existence of classical solutions to this nonlinear HJB equation for the value function. We then establish a verification theorem and analyze the policy-improvement structure induced by the entropy-regularized Hamiltonian, showing how the resulting sequence of PDEs provides a continuous-time interpretation of actor-critic learning dynamics. Finally, our PDE analysis with a semi-closed form of optimal value and optimal policy enables the design of an implementable reinforcement learning algorithm by recasting the optimal problem in a martingale framework.