Score
Designs and implements closed‑loop correction mechanisms that compute corrective control inputs by differentiating task error through system dynamics or differentiable operators, or alternatively by estimating corrections via zero‑order/derivative‑free queries when gradients are unavailable. Builds and analyzes algorithms that backpropagate task deviations into actionable changes (e.g., velocities), evaluates their stability and convergence, and trades off accuracy, robustness, and computational cost between gradient‑based and zero‑order strategies.
Robot path planning and trajectory optimization are commonly formulated as optimal control problems (OCPs), yet designing appropriate trade-offs among multi-objective cost components remains challenging, and resulting solutions often lack interpretability—leading to inefficient debugging. Method: We propose the first direction-corrected cost consistency analysis framework, integrating sensitivity analysis, gradient direction projection, and expert-feedback-driven iterative reweighting optimization. Contribution/Results: This approach enables interpretable diagnostic analysis of cost components and automated weight tuning, shifting from conventional trial-and-error to goal-directed correction. It significantly improves solution rationality and task success rates while supporting adaptive objective function reconstruction with low cost and minimal samples.
This study addresses the problem of neural trajectory predictors violating physical constraints when control inputs are unobserved, and proposes MaDE, a post-processing operator that projects state transitions onto a feasible manifold. By inferring unknown controls and rectifying inequality constraints, MaDE ensures dynamical consistency. Its key innovations include training without ground-truth control signals, guaranteeing iteratively optimal feasibility via gradient-based correction, and serving as a plug-and-play module for arbitrary prediction models. Experiments demonstrate that MaDE drives dynamical residuals to near zero in simulation and reduces them to 0.0071 on real-world vehicle data, outperforming baselines. This strict physical compliance is achieved at the cost of only an approximately 1.8-fold increase in displacement error.
This work addresses discrete-time nonlinear optimal control problems by unifying classical algorithms—including gradient descent, Gauss–Newton, Newton’s method, and differential dynamic programming (DDP)—within a differentiable programming framework. Methodologically, it introduces the first modular, end-to-end differentiable algorithm template library built upon linear/quadratic approximations (e.g., LQR), enabled by automatic differentiation. Theoretically, it provides a unified derivation of computational complexity and sufficient optimality conditions across all methods. Practically, it incorporates adaptive line search and regularization strategies, and validates efficacy on benchmark tasks such as autonomous racing with a bicycle model. All implementations are open-sourced, demonstrating both efficient gradient propagation and strong generalization across diverse control problems.
This work addresses the problem of stabilizing nonlinear, partially observable systems via neural control. We propose an end-to-end learning framework intrinsically guaranteeing incremental stability. Methodologically, we integrate nonlinear Youla–Kučera parameterization with recurrent equilibrium networks (RENs), establishing—for the first time—a joint d-tube contraction and Lipschitz continuity certification mechanism. This enables a complete, unconstrained parameterization of all contraction-Lipschitz closed-loop controllers. Our key contributions are: (1) overcoming the triply coupled challenges of nonlinear dynamics, partial observability, and incremental stability (i.e., contraction plus Lipschitz robustness); (2) ensuring weak yet practically meaningful closed-loop stability *naturally*, without explicit stability constraints during optimization; and (3) demonstrating—through experiments—that the method efficiently learns controllers with rigorous stability certificates under low-sample regimes, economic reward settings, and model uncertainty, significantly improving both robustness and convergence.
This paper addresses the infinite-horizon optimal closed-loop control problem for nonlinear systems with unknown dynamics, aiming to minimize a given cost function from arbitrary initial states without relying on an explicit system model. We propose a data-driven policy optimization method that integrates the Koopman operator with an actor-critic framework: the Koopman operator enables model-free dynamical representation and differentiable cost gradient estimation, while a parameterized policy is updated via stochastic gradient descent. To our knowledge, this is the first model-free policy gradient method with theoretically guaranteed convergence. Experiments demonstrate stable convergence across multiple nonlinear systems, with control performance significantly surpassing standard model-free reinforcement learning algorithms and closely approaching the optimal benchmark achievable under full model knowledge.
This study addresses the fragility of hyperparameter tuning and policy instability in deep reinforcement learning caused by heuristic-dependent reward shaping. Through theoretical analysis and empirical experiments in robotic control, it systematically investigates how zeroth- and first-order information affects policy gradients in stable control scenarios. We rigorously prove that control tasks can be accomplished without first-order reward terms, and that introducing such terms significantly increases policy gradient sensitivity. Based on these findings, we propose a “zeroth-order completeness” principle for reward design. This work provides theoretically grounded yet practically actionable reward design guidelines for robotic reinforcement learning, effectively reducing tuning complexity and enhancing training stability.
This work addresses the challenge of verifying closed-loop safety in safety-critical systems—such as aerospace applications—where neural network controllers act as black boxes. The authors propose a sampling-free reachability analysis framework that embeds a trained neural network into the system dynamics to form an autonomous closed-loop system. By integrating high-order Taylor expansions, automated domain partitioning, and polynomial bounding techniques, the method rigorously confines the propagation of state uncertainties over event manifolds. This approach enables, for the first time, a global, sound, and efficient reachability analysis of neural network-controlled systems, yielding tight upper and lower bounds on controller outputs across large state spaces. These guaranteed bounds facilitate reliable safety certification and informed mission-level decision-making.
This work addresses the challenge in behavioral cloning for position-controlled robots, where training loss often fails to predict real-world performance due to the absence of non-asymptotic theory characterizing how controller gains affect closed-loop behavior. The paper establishes the first gain-dependent, non-asymptotic framework for closed-loop error dynamics, revealing how sub-Gaussian action errors propagate through a PD controller into position errors. It decomposes task failure probability into a gain-dependent exponentially amplified term, validation loss, and a generalization gap. By introducing a surrogate matrix and its scalar upper bound, the analysis rigorously characterizes error tightness across four canonical control parameter regimes, theoretically justifying why compliant, overdamped controllers enhance cloning success. Leveraging stochastic process analysis, sub-Gaussian modeling, and linear system theory, the study derives a closed-form expression for continuous-time steady-state variance, proving its strict monotonicity with respect to stiffness and damping within the stability region—a property preserved in discrete-time systems, with theoretical predictions closely matching empirical results.
This work addresses the challenges posed by singular configurations in inverse kinematics for serial manipulators—such as loss of task-space mobility, unbounded joint velocities, and solver divergence—by proposing a unified framework that integrates Jacobian regularization, Riemannian manipulability tracking, constrained optimization, and data-driven techniques. It establishes, for the first time, a systematic taxonomy bridging classical robust inverse kinematics and learning-based approaches, categorizing existing methods according to the geometric structures they preserve and the nature of their robustness guarantees, whether formal or empirical. Evaluation of twelve solvers on the Franka Panda platform demonstrates that purely learning-based methods exhibit high failure rates, whereas hybrid architectures employing classical methods for refinement achieve significantly higher success rates of 98.6%–100%, thereby validating the efficacy and superiority of the proposed framework.
This work addresses the challenge that online updates of neural network controllers can compromise closed-loop stability. To overcome this, the authors propose a stability-preserving online update mechanism that, for the first time, enables stable switching of nonlinear neural controllers. By modeling the controller as an ℓp-gain-bounded causal operator, they derive sufficient gain conditions and design two update strategies—time-scheduled and state-triggered—that decouple stability guarantees from controller optimization, thereby accommodating approximate or early-stopped training. Under time-varying references and disturbances, the method continuously improves control performance while rigorously preserving closed-loop ℓp stability across multiple updates, significantly outperforming both static and naive online baselines.