Wasserstein Policy Gradient for Entropy-Regularized Linear-Quadratic Control

πŸ“… 2026-08-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses policy optimization in entropy-regularized discounted linear-quadratic (LQ) control. By introducing optimal transport over the action space, the Wasserstein policy gradient is exactly reduced to a finite-dimensional ordinary differential equation (ODE) governing the feedback gain and action covariance. This ODE is shown to be globally well-posed and to converge exponentially from any feasible initial condition. Notably, the convergence rate remains uniformly bounded as the entropy temperature approaches zero, circumventing the exponentially decaying rates that plague conventional methods under low entropy. Integrating Wasserstein gradient flows, Bellman verification, and ODE analysis, the study provides an exact characterization of the optimal linear-Gaussian policy and establishes a global exponential convergence guarantee, substantially enhancing the stability and efficiency of policy optimization.
πŸ“ Abstract
Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space. We study entropy-regularized discounted linear-quadratic (LQ) control. A Bellman verification argument shows that the unrestricted problem has a linear-Gaussian optimal policy, and the discounted-occupancy-weighted statewise Wasserstein gradient is tangent to this policy class. WPG therefore reduces exactly to a finite-dimensional ODE for the feedback gain and action covariance. We prove that this ODE is globally well posed and converges exponentially from every admissible initialization. For each fixed LQ problem, the exponent has a positive limit as the entropy temperature tends to zero and contains no perturbative factor of the form $\exp(-c/Ο„)$, while retaining the usual dependence on the conditioning of the control problem.
Problem

Research questions and friction points this paper is trying to address.

Wasserstein policy gradient
entropy-regularized control
linear-quadratic control
policy convergence
discounted optimal control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Wasserstein policy gradient
entropy-regularized control
linear-quadratic control
exponential convergence
finite-dimensional ODE
πŸ”Ž Similar Papers
No similar papers found.