🤖 AI Summary
This work addresses the limitation of existing stochastic optimal control methods, which lack explicit constraints on the distribution of reference trajectories and thus struggle to balance task performance with behavioral fidelity. The paper introduces, for the first time, a KL-divergence regularization at the trajectory distribution level into the stochastic optimal control framework. By leveraging Girsanov’s theorem, this regularization is reformulated as a quadratic penalty on drift deviation, yielding a modified cost function that preserves the dynamic programming structure. The authors derive the corresponding Hamilton–Jacobi–Bellman (HJB) equation and optimal policy, and obtain a closed-form solution under the linear-quadratic (LQ) setting. Experiments demonstrate that the regularization parameter effectively trades off task performance against adherence to the reference trajectory, enabling the use of reference dynamics learned from offline data.
📝 Abstract
We introduce trajectory-regularized stochastic optimal control (TRSOC), which augments standard stochastic optimal control (SOC) with a Kullback--Leibler (KL) divergence between controlled and reference trajectory distributions. Using Girsanov's theorem, the trajectory KL reduces to a quadratic drift mismatch penalty, yielding a modified running cost that preserves the dynamic programming (DP) structure. We derive the corresponding Hamilton--Jacobi--Bellman (HJB) equation and characterize the optimal policy. In the linear-quadratic (LQ) setting, the formulation admits a closed-form solution with an augmented control cost. Experiments show that the regularization parameter induces a trade-off between performance-driven and reference-preserving behavior, including cases with reference dynamics learned from offline data.