Score
Designs and evaluates algorithms or learned controllers that choose variable time-step schedules for numerical integrators or simulators, including online, coupling-aware, reparameterized/time-warped, or RL-based policies. These methods select per-step sizes to trade off accuracy, stability and conserved‑quantity/error bounds against computational cost, often aiming to generalize across integrators and operate without expert tuning.
This work addresses the challenge of online parameter adaptation for robot controllers operating continuously along a single trajectory—where environment dynamics, policy structure, and optimization objectives are all unknown, time-varying, and lack state resets or trajectory segmentation. To this end, we propose M-GAPS, a model-based online policy optimization algorithm. M-GAPS innovatively integrates joint reparameterization of the state space and policy class to significantly improve the optimization landscape; it further combines geometric nonlinear controller modeling with efficient policy gradient estimation to achieve both high data efficiency and strong robustness against disturbances. Hardware experiments on quadrotor and Ackermann-steering ground vehicles demonstrate that M-GAPS converges faster and adapts more effectively than manually segmented baseline controllers. It enables real-time adaptation to severe dynamic disturbances—including strong wind gusts and sudden payload changes—thereby overcoming the dual limitations of conventional adaptive control (limited flexibility) and reinforcement learning (low sample efficiency).
Multirate problems with multiple time scales pose significant challenges in adaptively selecting time steps of varying rates while balancing accuracy and efficiency. Method: This paper introduces two novel multirate adaptive controllers and constructs the first fifth-order embedded multirate exponential Runge–Kutta (MERK) method. Built upon the embedded multirate infinitesimal (MRI) framework, the method employs an explicit MERK nesting structure to enable efficient local error estimation and stepsize control, supporting high-accuracy adaptive integration across arbitrarily many time scales. Contribution/Results: Theoretical analysis and benchmark tests demonstrate that the proposed method achieves high-order accuracy while substantially reducing computational cost. It exhibits superior flexibility and robustness compared to state-of-the-art multirate methods. Moreover, it provides a systematic, practical guideline for controller design and method selection in multirate time integration.
Controller transfer across dynamically similar yet geometrically scaled systems typically requires laborious, system-specific parameter retuning. Method: This paper proposes a zero-shot, parameter-free transfer method based on nondimensional model predictive control (NMPC). By constructing a nondimensional dynamical model, the approach rigorously preserves dynamic similarity and enables direct cross-scale transfer of closed-loop control performance. Furthermore, it integrates Bayesian optimization with reinforcement learning to jointly optimize controller hyperparameters across multiple scales. Contribution/Results: This work presents the first synergistic co-design of nondimensional MPC and automated hyperparameter optimization, substantially enhancing controller generalizability and deployment efficiency. The method is validated on swing-up control of inverted pendulums and autonomous racing car navigation—achieving seamless transfer without any manual tuning. Open-source implementation is provided.
This work addresses the long-standing open problem of global convergence for single-sample, single-timescale Actor-Critic algorithms in continuous state-action spaces. Using the linear quadratic regulator (LQR) as a canonical model, we establish the first global convergence guarantee to an ε-optimal policy. Our analysis integrates tools from control theory (exploiting the analytic structure of LQR), stochastic approximation, policy gradient estimation, and nonconvex optimization. We rigorously prove that the algorithm converges to an ε-optimal policy with sample complexity O(ε⁻²), matching the information-theoretic lower bound in order. This result breaks prior theoretical dependencies on discrete state-action spaces or two-timescale stepsize regimes. It provides the first tight convergence guarantee for widely deployed single-timescale Actor-Critic methods in continuous domains, thereby bridging a critical gap between theoretical analysis and practical reinforcement learning applications.
This work addresses discrete-time nonlinear optimal control problems by unifying classical algorithms—including gradient descent, Gauss–Newton, Newton’s method, and differential dynamic programming (DDP)—within a differentiable programming framework. Methodologically, it introduces the first modular, end-to-end differentiable algorithm template library built upon linear/quadratic approximations (e.g., LQR), enabled by automatic differentiation. Theoretically, it provides a unified derivation of computational complexity and sufficient optimality conditions across all methods. Practically, it incorporates adaptive line search and regularization strategies, and validates efficacy on benchmark tasks such as autonomous racing with a bicycle model. All implementations are open-sourced, demonstrating both efficient gradient propagation and strong generalization across diverse control problems.
This work addresses the high computational burden of traditional finite-horizon linear quadratic regulator (LQR) control, which requires repeated online solution of differential Riccati equations and thus struggles to meet real-time demands. The paper introduces Deep Operator Networks (DeepONets) to the LQR problem for the first time, establishing an offline learning framework that directly maps time-varying system parameters to Riccati solution trajectories, thereby shifting the computational load from online execution to a one-time training phase. The proposed approach features a scalable network architecture tailored for matrix-valued time-varying signals and a progressive training strategy, accompanied by theoretical guarantees on error propagation of the approximate solution and closed-loop stability. Experiments demonstrate that the method achieves high accuracy, strong generalization, and substantial computational speedup across diverse time-varying and time-invariant LQR tasks, making it well-suited for parametric and real-time optimal control applications.
This work addresses the challenge of controlling unknown, rapidly evolving time-varying systems, which existing learning-based model predictive control (MPC) approaches struggle to handle effectively. To this end, we propose the T2S-MPC framework, which enables accurate real-time planning by online adaptive learning of residual dynamics and fusing them with a nominal model. Our approach incorporates structured temporal embeddings to endow the model with time-awareness and employs a dual-timescale update mechanism that balances rapid adaptation with learning stability. Experimental results demonstrate that T2S-MPC significantly outperforms classical MPC, neural MPC, and its ablated variants in stabilizing and trajectory-tracking tasks for a 2D quadrotor under diverse time-varying disturbances, exhibiting superior control performance and robustness.
This work proposes a mesh-free, near-symplectic integrator based on a learned continuous Lagrangian for long-term stable prediction of dynamical systems governed by partial differential equations (PDEs). By minimizing the Euler–Lagrange residual over local space-time blocks, the method constructs an optimization-based integrator that decouples model error from integration error, circumventing the global coupling inherent in conventional time-stepping schemes. It naturally accommodates arbitrary boundary conditions without retraining. To the best of our knowledge, this is the first approach to directly employ a learned Lagrangian for PDE evolution. The method achieves accuracy comparable to classical symplectic integrators in benchmark problems such as the double pendulum and one- and two-dimensional wave equations, and successfully generalizes to scenarios with spatially varying dynamics and complex boundary conditions.
This study addresses the lack of a systematic synthesis in research on integrating reinforcement learning (RL) with model predictive control (MPC) for linear systems by proposing the first multidimensional taxonomy tailored to this domain. Drawing on a comprehensive literature review up to 2025, the work establishes a classification framework along five dimensions: RL role, algorithm type, MPC formulation, cost function structure, and application area, followed by an integrative cross-dimensional analysis. The study elucidates representative integration strategies, traces methodological evolution, and identifies key challenges—including computational burden, sample efficiency, robustness, and closed-loop guarantees—thereby offering a structured reference and practical guidance for both theoretical analysis and architectural design in RL–MPC systems.
Traditional controllers struggle to generalize across systems with varying orders and dynamic characteristics. This work proposes a universal learning-based controller that constructs a dynamic state-space representation using a masked attention mechanism, integrating system label encoding, multi-scale temporal processing, and a mixture-of-experts architecture to enable a single neural network to uniformly control diverse linear and nonlinear systems. Notably, the approach is the first to adapt—without architectural modifications—to challenging dynamics such as unstable and non-minimum-phase systems, while supporting zero-shot generalization to unseen operating conditions. Trained on 25 system classes and 314,630 trajectories, the controller matches the performance of specialized LQI controllers and maintains robustness under previously unobserved conditions, including actuator saturation, noise, and disturbances.