Linear-Quadratic Mean-Field Reinforcement Learning: Convergence of Policy Gradient Methods

📅 2019-10-09
🏛️ arXiv.org
📈 Citations: 61
✨ Influential: 4
📄 PDF
🤖 AI Summary
This paper addresses the cooperative control of large-scale homogeneous agents in mean-field Markov decision processes (MFG-MDPs), focusing on optimizing a single-agent policy via reinforcement learning to minimize the aggregate social cost. For the mean-field linear-quadratic (MF-LQ) setting, we establish, for the first time, the global convergence of both exact and model-free policy gradient algorithms—providing the first theoretical convergence guarantee for mean-field reinforcement learning. Our approach integrates tools from linear-quadratic stochastic control, mean-field game modeling, and stochastic approximation theory, enabling convergence without prior knowledge of the environment dynamics. Numerical experiments confirm the predicted convergence rates and theoretical consistency. The key contribution is the first model-free policy gradient framework for distributed learning in large-scale agent systems with provable global convergence.
📝 Abstract
We investigate reinforcement learning for mean field control problems in discrete time, which can be viewed as Markov decision processes for a large number of exchangeable agents interacting in a mean field manner. Such problems arise, for instance when a large number of robots communicate through a central unit dispatching the optimal policy computed by minimizing the overall social cost. An approximate solution is obtained by learning the optimal policy of a generic agent interacting with the statistical distribution of the states of the other agents. We prove rigorously the convergence of exact and model-free policy gradient methods in a mean-field linear-quadratic setting. We also provide graphical evidence of the convergence based on implementations of our algorithms.
Problem

Research questions and friction points this paper is trying to address.

Study reinforcement learning for many exchangeable agents in mean-field interactions
Learn optimal policy for generic agent via state-action distribution of others
Prove convergence of policy gradient methods in linear-quadratic mean-field setting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mean-field reinforcement learning for exchangeable agents
Policy gradient methods convergence analysis
Linear-quadratic optimal policy approximation
🔎 Similar Papers
Princeton University | NYU Shanghai | ECNU | Walmart Global Tech
R
R. Carmona
Department of Operations Research and Financial Engineering & Program in Applied and Computational Mathematics, Princeton NJ 08544, USA
M
M. Laurière
Shanghai Frontiers Science Center of Artificial Intelligence and Deep Learning; NYU-ECNU Institute of Mathematical Sciences at NYU Shanghai; NYU Shanghai, 567 West Yangsi Road, Shanghai, 200126, People’s Republic of China
Z
Zongjun Tan
Department of Operations Research and Financial Engineering & Program in Applied and Computational Mathematics, Princeton NJ 08544, USA; Walmart Global Tech, USA