Score
Designs and analyzes policy-gradient algorithms for mean-field (large-population) control problems, deriving gradient expressions for feedback policies using tools such as the policy score or an infinitesimal-advantage representation and accommodating entropy-regularized randomized feedback policies. Implements and evaluates actor-critic and other model-free policy-gradient variants (including mean-field actor-critic and continuous-action representative methods such as DDPG-style updates) and establishes or empirically studies their convergence and performance properties.
This work addresses average-field games (MFGs), mean-field control (MFC), and hybrid MFC–game (MFCG) problems in continuous state spaces over infinite horizons. We propose the first unified actor–critic framework for solving all three problem classes. Our method parameterizes the score function—i.e., the gradient of the log-density—to implicitly represent the mean-field distribution, and couples it with Langevin dynamics for online sampling and distributional updates. Crucially, the solution objective—whether an MFG equilibrium, an MFC optimal policy, or a hybrid MFCG solution—is adaptively selected solely by tuning the learning rate. By bypassing explicit density modeling, our approach improves accuracy in distribution evolution and enhances policy–distribution coordination efficiency. On linear–quadratic benchmarks, the algorithm demonstrates stable convergence; both theoretical convergence guarantees and numerical robustness are rigorously validated.
This paper addresses the cooperative control of large-scale homogeneous agents in mean-field Markov decision processes (MFG-MDPs), focusing on optimizing a single-agent policy via reinforcement learning to minimize the aggregate social cost. For the mean-field linear-quadratic (MF-LQ) setting, we establish, for the first time, the global convergence of both exact and model-free policy gradient algorithms—providing the first theoretical convergence guarantee for mean-field reinforcement learning. Our approach integrates tools from linear-quadratic stochastic control, mean-field game modeling, and stochastic approximation theory, enabling convergence without prior knowledge of the environment dynamics. Numerical experiments confirm the predicted convergence rates and theoretical consistency. The key contribution is the first model-free policy gradient framework for distributed learning in large-scale agent systems with provable global convergence.
This paper addresses average-reward reinforcement learning for countable-state Markov decision processes (MDPs) with potentially unstable policies, where the stationary distribution belongs to an exponential family parameterized by the policy. Method: We propose the Score-Aware Gradient Estimator (SAGE), a value-function-free policy gradient estimator that directly exploits the exponential-family structure of the stationary distribution—bypassing the conventional actor-critic reliance on value function approximation. Contribution/Results: Theoretically, under non-convexity and infinite state spaces, we establish convergence guarantees via local Lyapunov conditions and Hessian non-degeneracy. Empirically, on multi-class product-form stochastic networks and queueing systems, SAGE achieves significantly faster training convergence to near-optimal policies compared to standard actor-critic methods, thereby validating both theoretical soundness and practical efficacy.
This paper addresses off-policy soft-maximum Actor-Critic algorithms under state distribution mismatch, where conventional density-ratio correction is infeasible or impractical. Method: We propose a unified finite-sample analysis framework for stochastic approximation algorithms operating on time-varying Markov chains, employing softmax policy parameterization, single-step stochastic updates, and an inexact critic—without requiring density-ratio correction, stationarity assumptions, or exact gradient access. Contribution/Results: For tabular MDPs, we establish the first global optimality guarantee for such off-policy Actor-Critic methods under these weak conditions. Our novel uniform contraction analysis tool enables rigorous finite-sample convergence characterization, yielding an $O(1/sqrt{T})$ rate. This significantly relaxes classical strong assumptions (e.g., ergodicity, exact gradients, or importance sampling), thereby enhancing theoretical interpretability and practical relevance to deep reinforcement learning training dynamics.
This work addresses the inverse problem of mean-field games (MFGs): reconstructing the obstacle function from partially observed value functions. To overcome the computational intractability of solving the coupled nonlinear forward–backward PDE system inherent in the forward MFG, we propose— for the first time—a decoupled framework based on policy iteration. Our method alternates between solving a linear PDE and a regularized linear inverse problem, inheriting the semantic structure of fixed-point iteration and provably achieving linear convergence. Numerical experiments in 1D and 2D using finite-difference discretization demonstrate that our approach significantly outperforms direct least-squares methods in accuracy, computational efficiency, robustness to observation noise, and scalability. To the best of our knowledge, this is the first inverse MFG solver that simultaneously offers rigorous theoretical guarantees and practical efficacy.
This work proposes a model-free reinforcement learning approach for continuous-time extended mean-field control problems where the joint distribution of states and controls exhibits explicit dependence. By adopting deterministic feedback policies, the state-action distribution is treated as the pushforward of the state distribution, thereby circumventing optimization over stochastic kernels. The study derives, for the first time, a policy gradient formula in Wasserstein space that incorporates both action derivatives and measure derivatives with respect to the control distribution. A martingale-driven learning framework is developed, integrating particle approximation, measure-dependent neural networks, temporal-difference learning, and exploration mechanisms to yield a continuous-time deep deterministic policy gradient algorithm tailored to this class of problems. The method demonstrates superior efficiency, stability, and robustness in applications including Cucker–Smale consensus control and optimal liquidation under trade crowding.
This study addresses the limitation of the standard REINFORCE algorithm in capturing population distribution effects within mean-field control by proposing the Transport REINFORCE method. This approach integrates model-free policy gradients, optimal transport mapping theory, and Gaussian mixture projection techniques. By perturbing transformed distributions on the probability simplex or Gaussian mixture manifold, it effectively estimates the missing mean-field contributions and policy gradients, accommodating both finite and continuous state spaces. Theoretically, we establish the consistency of the proposed estimator and derive its error bounds. Empirically, experiments demonstrate that Transport REINFORCE significantly outperforms the standard REINFORCE baseline, validating its effectiveness for mean-field control problems.
This work addresses the computational intractability arising from complex agent interactions in large-scale multi-agent reinforcement learning by proposing a scalable learning framework grounded in mean-field control theory. By leveraging mean-field approximation to characterize population behavior, the approach constructs a representative agent model and integrates it with a Markov decision process subject to common noise, thereby establishing a rigorous theoretical link between finite-population systems and their mean-field limits. The study presents the first systematic unification of mean-field control and reinforcement learning, offering formal analyses of propagation of chaos and algorithmic convergence. It further incorporates dynamic programming, Q-learning, policy gradient, and DDPG methods within this framework. Empirical validation on both general and linear-quadratic models demonstrates the efficacy of the proposed algorithms in efficiently approximating solutions for large-scale stochastic multi-agent systems.
This study addresses the mesh dependence and high-dimensional bottlenecks of Langevin policy iteration in entropy-regularized infinite-horizon stochastic control by proposing a mesh-free amortized actor-critic flow method. The approach projects the velocity field onto a shared conditional sampler and a parameterized critic, integrating a score-matching loss with an exact decomposition mechanism for Hamilton–Jacobi–Bellman residuals derived from the Feynman–Kac formula. This design enables efficient coupled optimization and establishes rigorous suboptimality bounds. The proposed method recovers pointwise iterative accuracy on linear-quadratic benchmarks and demonstrates its effectiveness across general high-dimensional models.
This work addresses continuous-time mean-field control problems where only discrete-time transition data are available, proposing a model-free reinforcement learning approach. By embedding the discrete observations into the continuous-time Hamilton–Jacobi–Bellman (HJB) equation over Wasserstein space, the method preserves the generator structure and circumvents identifiability issues. Building upon this formulation, the authors derive one-step estimation-based policy evaluation and policy gradient theorems, leading to an Actor-Critic algorithm termed MF-PhiBE. Theoretically, the value function approximation error is shown to be of order Δt; notably, in the linear-quadratic setting, second-order accuracy is achieved using only single-step data. Empirical validation on both linear-quadratic regulator (LQR) and crowd-avoidance tasks demonstrates the efficacy of the proposed method.