Score
Designs and implements iterative control and decision systems that use real-time measurements, solver outputs, and model predictions to update allocations or actions so as to optimize a specified objective under operational constraints. This work builds feedback-loop integration and model-in-the-loop validation to detect and correct model mismatch and disturbances, validate proposals with optimization, and iteratively reduce the optimality gap.
This work addresses discrete-time nonlinear optimal control problems by unifying classical algorithms—including gradient descent, Gauss–Newton, Newton’s method, and differential dynamic programming (DDP)—within a differentiable programming framework. Methodologically, it introduces the first modular, end-to-end differentiable algorithm template library built upon linear/quadratic approximations (e.g., LQR), enabled by automatic differentiation. Theoretically, it provides a unified derivation of computational complexity and sufficient optimality conditions across all methods. Practically, it incorporates adaptive line search and regularization strategies, and validates efficacy on benchmark tasks such as autonomous racing with a bicycle model. All implementations are open-sourced, demonstrating both efficient gradient propagation and strong generalization across diverse control problems.
To address the challenge of joint optimization between constraint embedding and policy learning in constrained reinforcement learning (CRL), this paper establishes a unified equivalence framework bridging CRL and feedback control. Specifically, Lagrange multiplier updates are reformulated as an optimal feedback control problem, and a multiplier-guided policy learning mechanism is introduced to enable end-to-end co-optimization. Theoretically, we show that PID-Lagrangian methods constitute only a special case within this broader framework. Methodologically, we pioneer the integration of model predictive control (MPC) into Lagrangian optimization, proposing Predictive Lagrangian Optimization (PLO)—a novel paradigm for adaptive constraint handling. Evaluated on a multi-task constrained RL benchmark, PLO significantly expands the feasible policy region (+7.2%) while preserving average reward performance, demonstrating its effectiveness, generalizability, and robustness.
In human-robot collaborative systems, real-time optimization of dynamic systems often becomes infeasible due to multiple conflicting soft constraints, posing safety risks. To address this, we propose a heuristic constraint selection mechanism leveraging historical Lagrange multipliers: at each optimization step, the least critical constraint is dynamically identified and removed to jointly ensure feasibility and safety. Our method integrates soft-constraint relaxation modeling, sensitivity analysis of Lagrange multipliers, heuristic constraint screening, and real-time convex optimization. Experiments demonstrate that the proposed approach matches state-of-the-art methods in control performance while reducing optimization variable dimensionality by 30–50% and accelerating per-step solving speed by 2.1×. The core innovation lies in the first use of historical multiplier evolution for online constraint importance assessment, enabling safety-driven, adaptive feasibility recovery.
For real-time parametric optimization problems (e.g., model predictive control), this paper proposes an end-to-end self-supervised neural iterative solver: a neural network first generates high-quality initial points, which are then refined by a differentiable primal-dual iterative module. The key contributions are twofold: (i) the design of the first KKT-based, label-free loss function, whose global minima are theoretically guaranteed to coincide exactly with KKT points; and (ii) a local convexification approximation strategy for non-convex problems, extending convergence guarantees to non-convex settings. The method requires no ground-truth labels and enables purely self-supervised training. Evaluated on two canonical non-convex benchmark tasks, it achieves a 10× speedup over IPOPT while attaining solution accuracy orders of magnitude higher than existing learning-based approaches.
This work addresses nonlinear systems subject to unknown dynamics and external disturbances. Methodologically, it proposes an integrated online system identification and model predictive control (MPC) framework that combines reproducing kernel Hilbert space (RKHS) modeling, random Fourier feature approximation, online least-squares parameter adaptation, and learning-based receding-horizon MPC—compatible with control-affine structures. The approach achieves sublinear dynamic regret against an adversarial clairvoyant controller for the first time, while ensuring finite-time near-optimality and asymptotic convergence to optimality. To jointly handle modeling errors and exogenous disturbances, it introduces self-supervised learning and state- and input-adaptive disturbance modeling. Extensive validation is conducted on an inverted pendulum, quadrotor simulation, and real-world quadrotor hardware under challenging conditions—including wind gusts, ground effect, and aerodynamic drag—demonstrating robustness and high-precision trajectory tracking performance.
This study addresses the lack of a systematic synthesis in research on integrating reinforcement learning (RL) with model predictive control (MPC) for linear systems by proposing the first multidimensional taxonomy tailored to this domain. Drawing on a comprehensive literature review up to 2025, the work establishes a classification framework along five dimensions: RL role, algorithm type, MPC formulation, cost function structure, and application area, followed by an integrative cross-dimensional analysis. The study elucidates representative integration strategies, traces methodological evolution, and identifies key challenges—including computational burden, sample efficiency, robustness, and closed-loop guarantees—thereby offering a structured reference and practical guidance for both theoretical analysis and architectural design in RL–MPC systems.
This work proposes a multi-agent collaborative framework that automatically translates natural language descriptions of operations research problems into solvable mathematical models and executable code. To address common modeling challenges—such as semantic misinterpretation, structural flaws, and mathematical inconsistencies—the approach employs specialized agents to extract decision variables and constraints, integrating structured information extraction, iterative self-correction, and a fourfold feedback validation mechanism to achieve end-to-end modeling. Its modular architecture enhances transparency and auditability throughout the modeling process. Evaluated on four standard benchmarks encompassing linear programming (LP), mixed-integer linear programming (MILP), and nonlinear programming, the method achieves state-of-the-art performance on three and demonstrates highly competitive results on the fourth.
This work proposes an information-theoretic iterative learning model predictive control framework to address the integrated demands of safety, robustness, and high performance in robotic iterative tasks operating under complex and uncertain environments. The approach leverages historical trajectories to learn a value function, employs normalizing flows to model non-Gaussian uncertainties, and incorporates an adaptive safety penalty mechanism to handle infinite-horizon constrained optimization for nonlinear stochastic systems. Efficient real-time solutions are achieved through highly parallelized GPU implementation. Both simulation and hardware experiments demonstrate that the system consistently improves control performance across iterations while rigorously satisfying safety constraints, thereby validating the effectiveness and superiority of the proposed method.
本文针对非线性系统中随机间歇测量的问题,提出了一种基于Koopman的鲁棒模型预测控制框架,并使用概率截断软约束方法来解决。
This work addresses the sensitivity to parameters, poor stability, and low computational efficiency commonly observed in gradient-based optimization algorithms for nonlinear model predictive control (NMPC). To overcome these limitations, the authors propose the Search and Accelerate (SaA) algorithm, which integrates adaptive line search, a trust-region mechanism, and gradient acceleration strategies. Specifically designed for box-constrained optimization problems, SaA requires no prior knowledge of the Lipschitz constant and operates effectively with default parameter settings, ensuring broad applicability. Theoretical analysis establishes its convergence and stability properties, while extensive experiments across 600 benchmark instances demonstrate its superior efficiency and robustness. Notably, SaA significantly reduces the NMPC control update cycle, positioning it as a compelling general-purpose alternative to existing gradient-based optimizers.