Wasserstein Policy Gradient for Entropy-Regularized Linear-Quadratic Control
2026-08-07 • Machine Learning
Machine Learning
AI summaryⓘ
The authors study a type of control problem where decisions are made continuously and influenced by some randomness, called entropy-regularized linear-quadratic control. They show that the best strategy in this setting can be described by a simple formula combining a linear part and random noise. By using a special kind of gradient called the Wasserstein policy gradient, they reduce the problem to solving a well-behaved ordinary differential equation (ODE). They prove this ODE converges quickly to the optimal strategy from any starting point, even as the randomness goes to zero, without losing stability.
Wasserstein policy gradiententropy regularizationlinear-quadratic controllinear-Gaussian policyBellman equationordinary differential equationdiscounted controloptimal controlconvergence analysispolicy gradient methods
Authors
Zhaoyu Zhu, Rui Gao, Shuang Li
Abstract
Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space. We study entropy-regularized discounted linear-quadratic (LQ) control. A Bellman verification argument shows that the unrestricted problem has a linear-Gaussian optimal policy, and the discounted-occupancy-weighted statewise Wasserstein gradient is tangent to this policy class. WPG therefore reduces exactly to a finite-dimensional ODE for the feedback gain and action covariance. We prove that this ODE is globally well posed and converges exponentially from every admissible initialization. For each fixed LQ problem, the exponent has a positive limit as the entropy temperature tends to zero and contains no perturbative factor of the form $\exp(-c/τ)$, while retaining the usual dependence on the conditioning of the control problem.