Score
Designs and implements algorithms and software that compute optimal control policies for dynamical systems, including formulating cost functionals, deriving necessary conditions, discretizing system dynamics, and building numerical solvers for trajectory optimization, feedback control, and model predictive control under state and input constraints. Analyzes and validates the correctness, convergence, and real-time performance of these implementations and the associated optimal control methods across continuous- and discrete-time problems.
This paper addresses the infinite-horizon optimal closed-loop control problem for nonlinear systems with unknown dynamics, aiming to minimize a given cost function from arbitrary initial states without relying on an explicit system model. We propose a data-driven policy optimization method that integrates the Koopman operator with an actor-critic framework: the Koopman operator enables model-free dynamical representation and differentiable cost gradient estimation, while a parameterized policy is updated via stochastic gradient descent. To our knowledge, this is the first model-free policy gradient method with theoretically guaranteed convergence. Experiments demonstrate stable convergence across multiple nonlinear systems, with control performance significantly surpassing standard model-free reinforcement learning algorithms and closely approaching the optimal benchmark achievable under full model knowledge.
This work addresses discrete-time nonlinear optimal control problems by unifying classical algorithms—including gradient descent, Gauss–Newton, Newton’s method, and differential dynamic programming (DDP)—within a differentiable programming framework. Methodologically, it introduces the first modular, end-to-end differentiable algorithm template library built upon linear/quadratic approximations (e.g., LQR), enabled by automatic differentiation. Theoretically, it provides a unified derivation of computational complexity and sufficient optimality conditions across all methods. Practically, it incorporates adaptive line search and regularization strategies, and validates efficacy on benchmark tasks such as autonomous racing with a bicycle model. All implementations are open-sourced, demonstrating both efficient gradient propagation and strong generalization across diverse control problems.
This work proposes a hybrid time-domain model predictive control (MPC) framework for hybrid dynamical systems subject to both continuous and discrete dynamics as well as state-input constraints. By integrating hybrid equation modeling, control Lyapunov functions, and terminal cost design, the approach formulates a structurally optimized MPC problem and establishes verifiable conditions for asymptotic stability. The method coherently links stage cost, terminal cost, and a static state feedback law, thereby theoretically guaranteeing asymptotic stability of the closed-loop system with respect to a target set. The efficacy and broad applicability of the proposed framework are demonstrated through multiple representative case studies.
This work addresses real-time trajectory optimization and cooperative control of autonomous agents on resource-constrained edge devices. We propose an efficient Model Predictive Control (MPC) framework based on integral Chebyshev collocation. Our key contribution is the first integration of integral Chebyshev polynomial parameterization with differentiable polyhedral collision checking, enabling explicit modeling of actuator saturation and hard obstacle-avoidance constraints. The formulation minimizes L₂ approximation error and is solved via quadratic programming for rapid convergence. Employing a receding-horizon MPC architecture, the method achieves over 3.2× speedup over conventional approaches on edge hardware, enabling sub-millisecond replanning. We validate its safety, real-time performance, and cooperative capability in multi-spacecraft formation control tasks, demonstrating robust constraint satisfaction and scalable coordination under tight computational budgets.
This work addresses the problem of designing data-driven controllers for nonlinear dynamical systems with verifiable closed-loop stability guarantees. Methodologically, it introduces a Koopman operator-based control framework that embeds stability certificates directly into the Extended Dynamic Mode Decomposition (EDMD) modeling process—specifically, by enforcing a vanishing proportional error bound at the origin and jointly optimizing the controller and a Lyapunov function via semidefinite programming (SDP), grounded in Lyapunov stability theory. The key contributions are: (i) a stability-oriented EDMD paradigm that enables control-objective-driven, quantifiable characterization of model approximation error; and (ii) an end-to-end stability certification mechanism. Experimental evaluation on multiple benchmark nonlinear systems demonstrates that the proposed approach achieves strict closed-loop stability while significantly improving control accuracy and robustness compared to existing methods.
This study addresses the challenge of efficiently solving the Hamilton–Jacobi–Bellman equation for optimal control of high-dimensional nonlinear control-affine systems. The authors propose a novel approach that leverages the Pontryagin maximum principle to generate training data comprising the value function, its gradient, and Hessian. By integrating hyperbolic cross sparse polynomial expansions with weighted least squares regression, they construct a high-fidelity approximation model. A key innovation lies in explicitly incorporating Hessian information into the supervised learning framework, complemented by a partial Hessian strategy that balances computational efficiency and approximation accuracy. Experimental results demonstrate that, in high-dimensional settings, the proposed method reduces the required number of training samples by nearly an order of magnitude compared to approaches using only function values, while significantly improving both value function approximation accuracy and closed-loop control performance.
This work addresses the sensitivity to parameters, poor stability, and low computational efficiency commonly observed in gradient-based optimization algorithms for nonlinear model predictive control (NMPC). To overcome these limitations, the authors propose the Search and Accelerate (SaA) algorithm, which integrates adaptive line search, a trust-region mechanism, and gradient acceleration strategies. Specifically designed for box-constrained optimization problems, SaA requires no prior knowledge of the Lipschitz constant and operates effectively with default parameter settings, ensuring broad applicability. Theoretical analysis establishes its convergence and stability properties, while extensive experiments across 600 benchmark instances demonstrate its superior efficiency and robustness. Notably, SaA significantly reduces the NMPC control update cycle, positioning it as a compelling general-purpose alternative to existing gradient-based optimizers.
This study addresses the computational bottleneck associated with dynamically constrained sampling of state spaces in feedback control and planning. To overcome this challenge, the work reformulates control as a dynamically constrained sampling problem, establishing a mapping framework that bridges controllability, optimal control theory, and generative modeling. Specifically, it integrates flow matching, normalizing flows, and denoising diffusion techniques to guide system evolution toward target states or distributions. The proposed approach enables efficient reachable set sampling and precise trajectory planning while unifying control-theoretic and generative-modeling paradigms. Furthermore, the authors provide an accessible open-source tutorial to facilitate practical adoption by the research community.
该研究通过引入差分博弈方法,解决单代理多目标动态系统中竞争控制任务的处理问题,提出了一种新的分而治之控制设计方法。
This work addresses the lack of global convergence guarantees in existing Differential Dynamic Programming (DDP) algorithms for optimal control problems with nonlinear state and control constraints. We propose FilterDDP, a novel algorithm that integrates a filter-based line search mechanism into the DDP framework. Instead of conventional damped Newton steps, FilterDDP generates search directions via backward recursion and trial points through forward simulation, thereby satisfying both dynamics and nonlinear constraints while ensuring iterative convergence. We provide the first rigorous proof of global convergence for this backward–forward procedure on a class of constrained optimal control problems and establish its theoretical equivalence to filter methods. This result offers a new solution paradigm that combines theoretical rigor with practical effectiveness for constrained optimal control.