mean-field policy gradient

Designs and analyzes policy-gradient algorithms for mean-field (large-population) control problems, deriving gradient expressions for feedback policies using tools such as the policy score or an infinitesimal-advantage representation and accommodating entropy-regularized randomized feedback policies. Implements and evaluates actor-critic and other model-free policy-gradient variants (including mean-field actor-critic and continuous-action representative methods such as DDPG-style updates) and establishes or empirically studies their convergence and performance properties.

mean-fieldpolicygradient

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Deep Reinforcement Learning for Infinite Horizon Mean Field Problems in Continuous Spaces

Sep 19, 2023
AA
Andrea Angiuli
🏛️ Amazon | University of California, Santa Barbara

This work addresses average-field games (MFGs), mean-field control (MFC), and hybrid MFC–game (MFCG) problems in continuous state spaces over infinite horizons. We propose the first unified actor–critic framework for solving all three problem classes. Our method parameterizes the score function—i.e., the gradient of the log-density—to implicitly represent the mean-field distribution, and couples it with Langevin dynamics for online sampling and distributional updates. Crucially, the solution objective—whether an MFG equilibrium, an MFC optimal policy, or a hybrid MFCG solution—is adaptively selected solely by tuning the learning rate. By bypassing explicit density modeling, our approach improves accuracy in distribution evolution and enhances policy–distribution coordination efficiency. On linear–quadratic benchmarks, the algorithm demonstrates stable convergence; both theoretical convergence guarantees and numerical robustness are rigorously validated.

Evaluates algorithm with linear-quadratic benchmarks in infinite horizon.Solves continuous-space mean field game and control problems.Uses actor-critic paradigm with parameterized score function.

Linear-Quadratic Mean-Field Reinforcement Learning: Convergence of Policy Gradient Methods

Oct 09, 2019
RC
R. Carmona
🏛️ Princeton University | NYU Shanghai | ECNU | Walmart Global Tech

This paper addresses the cooperative control of large-scale homogeneous agents in mean-field Markov decision processes (MFG-MDPs), focusing on optimizing a single-agent policy via reinforcement learning to minimize the aggregate social cost. For the mean-field linear-quadratic (MF-LQ) setting, we establish, for the first time, the global convergence of both exact and model-free policy gradient algorithms—providing the first theoretical convergence guarantee for mean-field reinforcement learning. Our approach integrates tools from linear-quadratic stochastic control, mean-field game modeling, and stochastic approximation theory, enabling convergence without prior knowledge of the environment dynamics. Numerical experiments confirm the predicted convergence rates and theoretical consistency. The key contribution is the first model-free policy gradient framework for distributed learning in large-scale agent systems with provable global convergence.

Learn optimal policy for generic agent via state-action distribution of othersProve convergence of policy gradient methods in linear-quadratic mean-field settingStudy reinforcement learning for many exchangeable agents in mean-field interactions

Score-Aware Policy-Gradient Methods and Performance Guarantees using Local Lyapunov Conditions: Applications to Product-Form Stochastic Networks and Queueing Systems

Dec 05, 2023
CC
Céline Comte
🏛️ CNRS | LAAS | Eindhoven University of Technology | IRIT | Université Toulouse III Paul Sabatier

This paper addresses average-reward reinforcement learning for countable-state Markov decision processes (MDPs) with potentially unstable policies, where the stationary distribution belongs to an exponential family parameterized by the policy. Method: We propose the Score-Aware Gradient Estimator (SAGE), a value-function-free policy gradient estimator that directly exploits the exponential-family structure of the stationary distribution—bypassing the conventional actor-critic reliance on value function approximation. Contribution/Results: Theoretically, under non-convexity and infinite state spaces, we establish convergence guarantees via local Lyapunov conditions and Hessian non-degeneracy. Empirically, on multi-class product-form stochastic networks and queueing systems, SAGE achieves significantly faster training convergence to near-optimal policies compared to standard actor-critic methods, thereby validating both theoretical soundness and practical efficacy.

Ensuring policy convergence using local Lyapunov stability analysisEstimating gradients without value-function approximation in MDPsImproving policy-gradient methods for model-based reinforcement learning

Global Optimality and Finite Sample Analysis of Softmax Off-Policy Actor Critic under State Distribution Mismatch

Nov 04, 2021
SZ
Shangtong Zhang
🏛️ University of Virginia | Microsoft Research Montreal

This paper addresses off-policy soft-maximum Actor-Critic algorithms under state distribution mismatch, where conventional density-ratio correction is infeasible or impractical. Method: We propose a unified finite-sample analysis framework for stochastic approximation algorithms operating on time-varying Markov chains, employing softmax policy parameterization, single-step stochastic updates, and an inexact critic—without requiring density-ratio correction, stationarity assumptions, or exact gradient access. Contribution/Results: For tabular MDPs, we establish the first global optimality guarantee for such off-policy Actor-Critic methods under these weak conditions. Our novel uniform contraction analysis tool enables rigorous finite-sample convergence characterization, yielding an $O(1/sqrt{T})$ rate. This significantly relaxes classical strong assumptions (e.g., ergodicity, exact gradients, or importance sampling), thereby enhancing theoretical interpretability and practical relevance to deep reinforcement learning training dynamics.

Analyze off-policy actor critic algorithm global optimalityConduct finite sample analysis with stochastic updatesRemove density ratio for state distribution correction

A Policy Iteration Method for Inverse Mean Field Games

Sep 10, 2024
KR
Kui Ren
🏛️ Columbia University

This work addresses the inverse problem of mean-field games (MFGs): reconstructing the obstacle function from partially observed value functions. To overcome the computational intractability of solving the coupled nonlinear forward–backward PDE system inherent in the forward MFG, we propose— for the first time—a decoupled framework based on policy iteration. Our method alternates between solving a linear PDE and a regularized linear inverse problem, inheriting the semantic structure of fixed-point iteration and provably achieving linear convergence. Numerical experiments in 1D and 2D using finite-difference discretization demonstrate that our approach significantly outperforms direct least-squares methods in accuracy, computational efficiency, robustness to observation noise, and scalability. To the best of our knowledge, this is the first inverse MFG solver that simultaneously offers rigorous theoretical guarantees and practical efficacy.

Decouple inverse MFG into linear PDEs and inverse problemsProve linear convergence and demonstrate superior efficiencyReconstruct obstacle function in MFG from partial observations

Latest Papers

What's happening recently
View more

This work proposes a model-free reinforcement learning approach for continuous-time extended mean-field control problems where the joint distribution of states and controls exhibits explicit dependence. By adopting deterministic feedback policies, the state-action distribution is treated as the pushforward of the state distribution, thereby circumventing optimization over stochastic kernels. The study derives, for the first time, a policy gradient formula in Wasserstein space that incorporates both action derivatives and measure derivatives with respect to the control distribution. A martingale-driven learning framework is developed, integrating particle approximation, measure-dependent neural networks, temporal-difference learning, and exploration mechanisms to yield a continuous-time deep deterministic policy gradient algorithm tailored to this class of problems. The method demonstrates superior efficiency, stability, and robustness in applications including Cucker–Smale consensus control and optimal liquidation under trade crowding.

deterministic policiesextended mean field controlMcKean–Vlasov dynamics

This study addresses the limitation of the standard REINFORCE algorithm in capturing population distribution effects within mean-field control by proposing the Transport REINFORCE method. This approach integrates model-free policy gradients, optimal transport mapping theory, and Gaussian mixture projection techniques. By perturbing transformed distributions on the probability simplex or Gaussian mixture manifold, it effectively estimates the missing mean-field contributions and policy gradients, accommodating both finite and continuous state spaces. Theoretically, we establish the consistency of the proposed estimator and derive its error bounds. Empirically, experiments demonstrate that Transport REINFORCE significantly outperforms the standard REINFORCE baseline, validating its effectiveness for mean-field control problems.

mean-field controlmodel-free reinforcement learningpolicy gradient

This work addresses the computational intractability arising from complex agent interactions in large-scale multi-agent reinforcement learning by proposing a scalable learning framework grounded in mean-field control theory. By leveraging mean-field approximation to characterize population behavior, the approach constructs a representative agent model and integrates it with a Markov decision process subject to common noise, thereby establishing a rigorous theoretical link between finite-population systems and their mean-field limits. The study presents the first systematic unification of mean-field control and reinforcement learning, offering formal analyses of propagation of chaos and algorithmic convergence. It further incorporates dynamic programming, Q-learning, policy gradient, and DDPG methods within this framework. Empirical validation on both general and linear-quadratic models demonstrates the efficacy of the proposed algorithms in efficiently approximating solutions for large-scale stochastic multi-agent systems.

common noiselarge-population stochastic controlMarkov decision processes

This study addresses the mesh dependence and high-dimensional bottlenecks of Langevin policy iteration in entropy-regularized infinite-horizon stochastic control by proposing a mesh-free amortized actor-critic flow method. The approach projects the velocity field onto a shared conditional sampler and a parameterized critic, integrating a score-matching loss with an exact decomposition mechanism for Hamilton–Jacobi–Bellman residuals derived from the Feynman–Kac formula. This design enables efficient coupled optimization and establishes rigorous suboptimality bounds. The proposed method recovers pointwise iterative accuracy on linear-quadratic benchmarks and demonstrates its effectiveness across general high-dimensional models.

entropy-regularized controlgrid-freeinfinite-horizon

This work addresses continuous-time mean-field control problems where only discrete-time transition data are available, proposing a model-free reinforcement learning approach. By embedding the discrete observations into the continuous-time Hamilton–Jacobi–Bellman (HJB) equation over Wasserstein space, the method preserves the generator structure and circumvents identifiability issues. Building upon this formulation, the authors derive one-step estimation-based policy evaluation and policy gradient theorems, leading to an Actor-Critic algorithm termed MF-PhiBE. Theoretically, the value function approximation error is shown to be of order Δt; notably, in the linear-quadratic setting, second-order accuracy is achieved using only single-step data. Empirical validation on both linear-quadratic regulator (LQR) and crowd-avoidance tasks demonstrates the efficacy of the proposed method.

continuous-time reinforcement learningdiscrete-time dataMcKean-Vlasov dynamics

Hot Scholars

HP

Haixia Peng

Professor, School of Information and Communications Engineering, Xi'an Jiaotong University
Resource managementReinforcement learningMulti-access edge computingSDN
CS

Clifford Stein

Professor of IEOR and CS, Columbia University
AlgorithmsOptimizationOperations ResearchScheduling
ZS

Zhou Su

Xi'an Jiaotong University
NC

Nan Cheng

University of Michigan
condensed matter physics
CZ

Conghao Zhou

School of Telecomm. Engineering, Xidian University
Immersive CommunicationAI for NetworkingNetwork Digital TwinSAGIN