Reinforcement Learning with Model Predictive Control for Highway Ramp Metering

📅 2023-11-15
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address congestion at highway on-ramp merging zones, this paper proposes a cooperative ramp metering method integrating reinforcement learning (RL) and model predictive control (MPC). The approach embeds policy-gradient RL algorithms—specifically PPO and SAC—into a differentiable MPC framework built upon the Cell Transmission Model (CTM) of traffic flow. A stage cost function is designed to jointly optimize traffic states, ensure control smoothness, and enforce hard queue-length constraints. Crucially, the MPC optimization problem itself serves as a differentiable functional approximator for the RL policy, enabling end-to-end policy optimization while guaranteeing strict constraint satisfaction. Evaluated on a benchmark highway network, the method significantly reduces mainline congestion, achieves 100% compliance with queue constraints, and outperforms conventional MPC and state-of-the-art adaptive controllers across all metrics. Notably, it demonstrates superior robustness and adaptability under model mismatch and dynamic traffic demand fluctuations.
📝 Abstract
In the backdrop of an increasingly pressing need for effective urban and highway transportation systems, this work explores the synergy between model-based and learning-based strategies to enhance traffic flow management by use of an innovative approach to the problem of ramp metering control that embeds Reinforcement Learning (RL) techniques within the Model Predictive Control (MPC) framework. The control problem is formulated as an RL task by crafting a suitable stage cost function that is representative of the traffic conditions, variability in the control action, and violations of the constraint on the maximum number of vehicles in queue. An MPC-based RL approach, which leverages the MPC optimal problem as a function approximation for the RL algorithm, is proposed to learn to efficiently control an on-ramp and satisfy its constraints despite uncertainties in the system model and variable demands. Simulations are performed on a benchmark small-scale highway network to compare the proposed methodology against other state-of-the-art control approaches. Results show that, starting from an MPC controller that has an imprecise model and is poorly tuned, the proposed methodology is able to effectively learn to improve the control policy such that congestion in the network is reduced and constraints are satisfied, yielding an improved performance that is superior to the other controllers.
Problem

Research questions and friction points this paper is trying to address.

Highway Ramp Management
Traffic Congestion
Traffic Flow Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Model Predictive Control
Traffic Flow Optimization
🔎 Similar Papers
No similar papers found.
Delft University of Technology
F
Filippo Airaldi
Delft Center for Systems and Control, Delft University of Technology, Mekelweg 2, 2628 CD Delft, The Netherlands
B
B. D. Schutter
Delft Center for Systems and Control, Delft University of Technology, Mekelweg 2, 2628 CD Delft, The Netherlands
A
Azita Dabiri
Delft Center for Systems and Control, Delft University of Technology, Mekelweg 2, 2628 CD Delft, The Netherlands