🤖 AI Summary
This study addresses the mesh dependence and high-dimensional bottlenecks of Langevin policy iteration in entropy-regularized infinite-horizon stochastic control by proposing a mesh-free amortized actor-critic flow method. The approach projects the velocity field onto a shared conditional sampler and a parameterized critic, integrating a score-matching loss with an exact decomposition mechanism for Hamilton–Jacobi–Bellman residuals derived from the Feynman–Kac formula. This design enables efficient coupled optimization and establishes rigorous suboptimality bounds. The proposed method recovers pointwise iterative accuracy on linear-quadratic benchmarks and demonstrates its effectiveness across general high-dimensional models.
📝 Abstract
We develop an amortized, grid-free implementation of continuous Langevin dynamics based policy-value iteration for entropy-regularized, infinite-horizon relaxed stochastic control problems. The improvement rate of the exact iteration is a discounted aggregate of relative Fisher information between the policy and the Gibbs law of its Hamiltonian. The associated score residual is the velocity with which the control's Langevin dynamics transport its law. We project this velocity onto a conditional sampler shared across states, instead of one Langevin dynamics per state, and the value dynamics onto a parametric critic, estimating both projections at sampled states to obtain coupled actor--critic flows. The score loss measures the actor's agreement with the current critic, while the policy-evaluation residual measures the critic's agreement with the actor. We also derive gradient and Hessian residuals, including a Feynman--Kac representation for the gradient equation, to control errors not detected by the projected value iteration. An exact decomposition of the HJB residual combines these errors into a policy-suboptimality bound under verification and logarithmic Sobolev assumptions. In the linear-quadratic class, both projections are exact and recover the pointwise iteration, and we provide numerical experiments on general models to demonstrate the coupled actor--critic learning in high-dimensions.