🤖 AI Summary
This study addresses the limitations of model predictive control (MPC), which relies on analytical models and struggles with complex state dependencies, as well as the low data efficiency inherent in reinforcement learning. To overcome these challenges, this work proposes a hybrid learning-based MPC framework that integrates local model-based planning with learned components. Specifically, the method employs a residual dynamics network to compensate for model bias and introduces an action-value function to incorporate long-horizon structural information. Furthermore, a GPU-accelerated batch iLQR solver is designed to enable efficient parallel computation. Experimental results demonstrate that the proposed framework significantly improves closed-loop control performance and data efficiency in tasks involving modeling discrepancies, while preserving the structured advantages characteristic of model-based approaches.
📝 Abstract
Model Predictive Control (MPC) provides a structured and constraint-aware mechanism for decision-making, but its reliance on optimization-friendly analytical dynamics models limits its use in tasks with contacts and other hard-to-model state dependencies. Model-free reinforcement learning avoids explicit modeling assumptions but typically requires large amounts of interaction data. We present a learning-based MPC framework that combines the data efficiency and structure of local model-based planning with learned components that compensate for incomplete dynamics and finite-horizon myopia. The method augments a nominal analytical model with a residual dynamics network that learns missing state-dependent effects from data and combines the resulting planner with a learned action-value critic that injects long-horizon MDP structure into the local iLQR optimization. To make this practical at reinforcement-learning scale, we develop a GPU-accelerated batched iLQR solver that evaluates learned dynamics and critic networks inside the optimal-control loop and solves thousands of trajectory-optimization problems in parallel. The complete system is integrated into a robotics simulator, enabling scalable model-based reinforcement learning under incomplete dynamics. Experiments on biased and incompletely modeled control tasks show that the approach improves closed-loop control performance while preserving the model-based structure needed for efficient constrained trajectory optimization.