task-adaptive mpc

Designs, implements, and analyzes receding‑horizon, optimization‑based controllers that adapt the prediction horizon and task‑specific objectives or active subtasks online; this includes formulating and re‑solving constrained optimization problems at each control step using the current measured state, selecting or switching active subtasks (for example via logical or automata rules), and truncating the prediction horizon to the remaining task window to reduce per‑step computation for closed‑loop control.

task-adaptivempc

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Suboptimality analysis of receding horizon quadratic control with unknown linear systems and its applications in learning-based control

Jan 19, 2023
SS
Shengli Shi
🏛️ Massachusetts Institute of Technology | Delft University of Technology | ETH Zurich

This paper investigates the suboptimality of nominal model-predictive linear-quadratic (LQ) control for unknown linear systems, characterizing a fundamental trade-off among model mismatch, terminal cost approximation error, and prediction horizon length. We develop a novel perturbation analysis framework for the Riccati difference equation, establishing—for the first time—a quantitative relationship between horizon length and the system’s controllability index. Theoretically, we prove that a finite horizon bounded by the controllability index suffices to approximate infinite-horizon optimal performance, and that horizons of length one or infinity are often optimal. Based on this insight, we derive the first adaptive horizon-selection criterion tailored for learning-based control, yielding a tight suboptimality upper bound, an $O(log T)$ regret guarantee, and optimal sample complexity.

Analyzes performance trade-offs in receding-horizon LQ control with modeling errors.Applies suboptimality bounds to learning-based control for regret guarantees.Determines optimal prediction horizon for near-optimal control performance.

This work addresses the challenge of real-time robotic arm control in dynamically cluttered environments, where agents must balance rapid responsiveness with foresightful obstacle avoidance to prevent myopic constraint violations. The authors propose a task-space receding horizon controller that generates collision-free terminal pose references through short-horizon, contact-consistent forward simulations respecting non-penetration constraints, then computes only the first-step minimum-acceleration control input that smoothly transitions toward this reference. By integrating the strengths of receding horizon and reactive control, the method efficiently embeds information about contacts, moving obstacles, and self-collisions using inflated convex geometry and an iterative dynamics solver—without requiring full trajectory optimization. Simulations with 40 degrees of freedom demonstrate that a moderate horizon length effectively balances foresight, responsiveness, and computational cost, while hardware experiments on a 6-DOF manipulator confirm strong sim-to-real transfer, outperforming MPC and dynamic optimization fabric approaches in success rate under dynamic clutter while meeting real-time requirements.

collision avoidancedynamic obstaclesreal-time control

Model Predictive Control (MPC) suffers from high computational complexity in long horizons and difficulty guaranteeing closed-loop stability in short horizons. To address this, we propose the “Observed Control” framework, which exploits the rigorous duality between state estimation and MPC. It employs the Kalman smoother as a unified optimization backbone to explicitly decouple the reactive and predictive components of the control law. The framework supports online control synthesis for arbitrary horizon lengths and incorporates an adaptive optimization termination criterion to enable early convergence. By integrating extended or unscented Kalman filtering for nonlinear systems, the method ensures closed-loop stability while reducing computational complexity to linear in horizon length. Numerical experiments on nonlinear systems demonstrate its efficiency, scalability, and robustness—achieving real-time performance without sacrificing stability or prediction fidelity.

Efficiently compute control actions with linear scalabilityEnsure controller stability for short time horizonsExtend observed control to nonlinear systems and objectives

This work proposes Drifting MPC, a novel framework for offline reinforcement learning in settings where the system dynamics are unknown and trajectory simulation is infeasible. Drifting MPC uniquely integrates a drift generative model with model predictive control to learn a conditional trajectory distribution from offline data that balances data support and cost optimality. The method explicitly optimizes task-specific costs while maintaining fidelity to the empirical data distribution, and it is theoretically shown that the resulting distribution constitutes the unique solution that optimally trades off optimality against consistency with the data prior. Empirical results demonstrate that Drifting MPC efficiently generates near-optimal trajectories with low per-step computational overhead, significantly reducing trajectory generation time compared to diffusion-model baselines.

generative modelsoffline datasetreceding-horizon control

Predictive Lagrangian Optimization for Constrained Reinforcement Learning

Jan 25, 2025
TZ
Tianqi Zhang
🏛️ Tsinghua University | University of Science and Technology Beijing

To address the challenge of joint optimization between constraint embedding and policy learning in constrained reinforcement learning (CRL), this paper establishes a unified equivalence framework bridging CRL and feedback control. Specifically, Lagrange multiplier updates are reformulated as an optimal feedback control problem, and a multiplier-guided policy learning mechanism is introduced to enable end-to-end co-optimization. Theoretically, we show that PID-Lagrangian methods constitute only a special case within this broader framework. Methodologically, we pioneer the integration of model predictive control (MPC) into Lagrangian optimization, proposing Predictive Lagrangian Optimization (PLO)—a novel paradigm for adaptive constraint handling. Evaluated on a multi-task constrained RL benchmark, PLO significantly expands the feasible policy region (+7.2%) while preserving average reward performance, demonstrating its effectiveness, generalizability, and robustness.

Complex Task SolvingConstrained Reinforcement LearningRule Integration

Latest Papers

What's happening recently
View more

This study addresses the lack of a systematic synthesis in research on integrating reinforcement learning (RL) with model predictive control (MPC) for linear systems by proposing the first multidimensional taxonomy tailored to this domain. Drawing on a comprehensive literature review up to 2025, the work establishes a classification framework along five dimensions: RL role, algorithm type, MPC formulation, cost function structure, and application area, followed by an integrative cross-dimensional analysis. The study elucidates representative integration strategies, traces methodological evolution, and identifies key challenges—including computational burden, sample efficiency, robustness, and closed-loop guarantees—thereby offering a structured reference and practical guidance for both theoretical analysis and architectural design in RL–MPC systems.

IntegrationLinear SystemsModel Predictive Control

This work addresses the sensitivity to parameters, poor stability, and low computational efficiency commonly observed in gradient-based optimization algorithms for nonlinear model predictive control (NMPC). To overcome these limitations, the authors propose the Search and Accelerate (SaA) algorithm, which integrates adaptive line search, a trust-region mechanism, and gradient acceleration strategies. Specifically designed for box-constrained optimization problems, SaA requires no prior knowledge of the Lipschitz constant and operates effectively with default parameter settings, ensuring broad applicability. Theoretical analysis establishes its convergence and stability properties, while extensive experiments across 600 benchmark instances demonstrate its superior efficiency and robustness. Notably, SaA significantly reduces the NMPC control update cycle, positioning it as a compelling general-purpose alternative to existing gradient-based optimizers.

algorithm stabilitybox-constrained optimizationgradient-based optimization

This work addresses the stochastic optimal control problem with joint chance constraints over an infinite horizon. By augmenting the state space, the problem is reformulated as a constrained Markov decision process with an additive structure. The paper establishes strong duality for this setting for the first time, thereby equivalently transforming the original problem into an unconstrained Lagrangian dual problem. Building on this duality result, the authors propose a hybrid solution framework that integrates dual ascent with offline value function approximation. This approach significantly reduces online computational complexity while preserving both optimality and probabilistic feasibility of the solution. Numerical experiments demonstrate that, compared to existing online model predictive control strategies, the proposed method achieves comparable control performance with substantially improved computational efficiency.

chance constraintsinfinite-horizonMarkov decision process

This study addresses the finite-horizon budget allocation problem under non-stationary changes in return efficiency by formulating it as a closed-loop economic control problem. The authors employ a receding-horizon model predictive control (MPC) approach to dynamically optimize budget allocation, accounting for execution noise and operational constraints. Through comparison with reactive strategies, the research demonstrates that non-stationarity alone is insufficient for MPC to outperform reactive methods; MPC achieves significant and sustained superiority only when the return efficiency exhibits predictable structures that the model can effectively capture, thereby enabling advantageous intertemporal trade-offs. In contrast, under scenarios of random drift or stationarity, MPC offers no notable performance advantage over reactive approaches.

budget allocationeconomic controlintertemporal trade-offs

This work addresses the problem of adaptive control for stochastic linear quadratic regulators (LQR) with time-varying chance constraints. The authors propose a safe, optimism-based exploration method formulated via semidefinite programming (SDP), which selects optimistic policies while progressively retracting to verifiably safe ones, thereby satisfying safety constraints at every step while achieving low regret. The key innovation lies in establishing, for the first time in constrained LQR settings, a regret bound of $\tilde{O}(\sqrt{T})$, improving upon the prior best-known rate of $\tilde{O}(T^{2/3})$. The approach handles unbounded process noise through chance constraints and introduces a novel analytical framework based on system covariance—replacing conventional cost-function-based analyses—to theoretically guarantee both safety and near-optimal performance.

adaptive controlchance constraintsconstrained LQR

Hot Scholars

SK

Sangbae Kim

Professor of Mechanical Engineering, MIT
Robotics
AW

Antonia Wachter-Zeh

Professor at Technical University of Munich (TUM)
coding theorycryptography
VW

Violetta Weger

Technical University of Munich
Applied AlgebraCryptographyNumber TheoryCoding Theory