Score
Formulates decision problems as optimal-control or calculus-of-variations problems and builds control-system representations (states, controls, dynamics, objective) to analyze them. Derives necessary and sufficient optimality conditions (Euler–Lagrange, Pontryagin/first-order conditions, transversality), computes optimal trajectories or control laws, and uses these conditions to prove policy optimality, rule out deviations, and characterize solution properties.
This paper investigates necessary conditions for optimal Markovian control policies in finite-horizon stochastic control problems. Departing from conventional approaches based on dynamic programming or the stochastic maximum principle, it systematically extends classical calculus of variations to stochastic settings—integrating Itô calculus and stochastic analysis to rigorously derive first-order variational conditions for optimality, namely a stochastic Euler–Lagrange-type equation. Crucially, this framework does not require differentiability of the value function, thereby providing a novel necessity analysis tool for nonlinear and nonconvex stochastic control problems. As a key validation, the paper solves the classical Merton portfolio optimization problem analytically; the resulting explicit optimal policy coincides exactly with established results, confirming both the mathematical rigor and practical efficacy of the proposed method.
This paper establishes the theoretical foundations of infinite-horizon linear dynamic optimization models. Addressing core issues—including existence of solutions, sufficiency of optimality conditions, and validity of the dynamic programming equation—the study employs convex analysis, infinite-dimensional optimization, and the theory of upper semicontinuous set-valued mappings. It provides the first rigorous proof that the transversality condition unconditionally ensures optimality in such models; reveals that optimal decision rules must be upper semicontinuous correspondences—not necessarily single-valued functions—in linear settings; introduces the novel “two-stage linear cake-eating problem” paradigm and derives necessary conditions for its solution; and, under convex bi-periodic constraints, establishes concavity, continuity, and monotonicity of the value function, along with conditional monotonicity of the policy correspondence. Collectively, these results unify and extend the applicability of the Euler equation, transversality condition, and dynamic programming principle to linear dynamic optimization.
This work addresses the common omission of fundamental physical constraints—such as inertia, gravity, and viscous drag—in robot motion optimization. We propose a unified optimal control framework grounded in differential geometry. Methodologically, we model viscous drag as a Riemannian metric on the configuration manifold, thereby unifying kinetic energy and gravitational potential fields, and derive geometric optimal control equations incorporating curvature effects. Indirect optimal control is solved via Lagrangian mechanics and manifold-based variational calculus. Experiments on a two-link planar manipulator and a UR5 robot demonstrate that the proposed model substantially alters optimal trajectory shape and energy distribution, enhancing both physical realizability and energy efficiency. Our core contribution is a novel geometric modeling paradigm that synergistically integrates drag, curvature, and potential fields—establishing a new theoretical foundation for physics-informed robotic trajectory optimization.
This work addresses discrete-time nonlinear optimal control problems by unifying classical algorithms—including gradient descent, Gauss–Newton, Newton’s method, and differential dynamic programming (DDP)—within a differentiable programming framework. Methodologically, it introduces the first modular, end-to-end differentiable algorithm template library built upon linear/quadratic approximations (e.g., LQR), enabled by automatic differentiation. Theoretically, it provides a unified derivation of computational complexity and sufficient optimality conditions across all methods. Practically, it incorporates adaptive line search and regularization strategies, and validates efficacy on benchmark tasks such as autonomous racing with a bicycle model. All implementations are open-sourced, demonstrating both efficient gradient propagation and strong generalization across diverse control problems.
This paper addresses the infinite-horizon optimal closed-loop control problem for nonlinear systems with unknown dynamics, aiming to minimize a given cost function from arbitrary initial states without relying on an explicit system model. We propose a data-driven policy optimization method that integrates the Koopman operator with an actor-critic framework: the Koopman operator enables model-free dynamical representation and differentiable cost gradient estimation, while a parameterized policy is updated via stochastic gradient descent. To our knowledge, this is the first model-free policy gradient method with theoretically guaranteed convergence. Experiments demonstrate stable convergence across multiple nonlinear systems, with control performance significantly surpassing standard model-free reinforcement learning algorithms and closely approaching the optimal benchmark achievable under full model knowledge.
This study addresses a linear-quadratic stochastic optimal control problem subject to state constraints, aiming to steer the system trajectory away from prescribed forbidden regions in space-time while minimizing the expected cost of state and control. By modeling the dynamics via diffusion processes and leveraging stochastic control theory together with probabilistic representation techniques, the authors establish a probabilistic representation of the value function under regularity conditions on the constraint set and derive its explicit solution. The resulting optimal control policy is strongly adapted and implementable via the filtration generated by the underlying Brownian motion. In addition to providing several analytical examples, this work offers a systematic framework for solving stochastic control problems with state constraints.
This study addresses the challenge of efficiently solving the Hamilton–Jacobi–Bellman equation for optimal control of high-dimensional nonlinear control-affine systems. The authors propose a novel approach that leverages the Pontryagin maximum principle to generate training data comprising the value function, its gradient, and Hessian. By integrating hyperbolic cross sparse polynomial expansions with weighted least squares regression, they construct a high-fidelity approximation model. A key innovation lies in explicitly incorporating Hessian information into the supervised learning framework, complemented by a partial Hessian strategy that balances computational efficiency and approximation accuracy. Experimental results demonstrate that, in high-dimensional settings, the proposed method reduces the required number of training samples by nearly an order of magnitude compared to approaches using only function values, while significantly improving both value function approximation accuracy and closed-loop control performance.
This work proposes a novel paradigm termed “Optimized Natural Physics” to investigate whether optimization algorithms adhere to natural laws of motion induced by the objective function. By establishing equivalence between optimal control problems and generalized KKT conditions, the authors construct a natural vector field governed by non-Newtonian dynamics. Leveraging Pontryagin’s minimum principle, Hamilton–Jacobi inequalities, and energy dissipation mechanisms, they design control strategies possessing inverse optimality. This framework not only unifies the interpretation of diverse existing optimization algorithms but also enables the systematic derivation of new ones. The approach demonstrates that global optimization can be achieved through deliberate modulation of jumps and dissipation, thereby providing a physically intuitive and mathematically unified foundation for optimization theory.
This work addresses the limitation of existing stochastic optimal control methods, which lack explicit constraints on the distribution of reference trajectories and thus struggle to balance task performance with behavioral fidelity. The paper introduces, for the first time, a KL-divergence regularization at the trajectory distribution level into the stochastic optimal control framework. By leveraging Girsanov’s theorem, this regularization is reformulated as a quadratic penalty on drift deviation, yielding a modified cost function that preserves the dynamic programming structure. The authors derive the corresponding Hamilton–Jacobi–Bellman (HJB) equation and optimal policy, and obtain a closed-form solution under the linear-quadratic (LQ) setting. Experiments demonstrate that the regularization parameter effectively trades off task performance against adherence to the reference trajectory, enabling the use of reference dynamics learned from offline data.