🤖 AI Summary
This paper addresses the cooperative control of large-scale homogeneous agents in mean-field Markov decision processes (MFG-MDPs), focusing on optimizing a single-agent policy via reinforcement learning to minimize the aggregate social cost. For the mean-field linear-quadratic (MF-LQ) setting, we establish, for the first time, the global convergence of both exact and model-free policy gradient algorithms—providing the first theoretical convergence guarantee for mean-field reinforcement learning. Our approach integrates tools from linear-quadratic stochastic control, mean-field game modeling, and stochastic approximation theory, enabling convergence without prior knowledge of the environment dynamics. Numerical experiments confirm the predicted convergence rates and theoretical consistency. The key contribution is the first model-free policy gradient framework for distributed learning in large-scale agent systems with provable global convergence.
📝 Abstract
We investigate reinforcement learning for mean field control problems in discrete time, which can be viewed as Markov decision processes for a large number of exchangeable agents interacting in a mean field manner. Such problems arise, for instance when a large number of robots communicate through a central unit dispatching the optimal policy computed by minimizing the overall social cost. An approximate solution is obtained by learning the optimal policy of a generic agent interacting with the statistical distribution of the states of the other agents. We prove rigorously the convergence of exact and model-free policy gradient methods in a mean-field linear-quadratic setting. We also provide graphical evidence of the convergence based on implementations of our algorithms.