Score
Designs and analyzes reduced, averaged descriptions of large interacting systems using statistical mechanics and mean-field methods, deriving mean-field limit equations, mean-field reductions, linearizations, and multi-population reduced models that eliminate microscopic variables and approximate interactions by their averaged effects. Uses those reduced descriptions to compute phase diagrams and order parameters, predict cascades and typical-case learning properties, quantify stability/fragility and response functions, and formulate mean-field control or mean-field reinforcement-learning analyses (including mean-field Q‑learning).
This work addresses the challenge of approximating the mean-field dynamics of interacting particle systems exhibiting collective behavior using Transformer models, under the constraint of permutation equivariance arising from particle indistinguishability. Method: We propose the first theoretical Transformer framework that maps the vector field of a finite-particle system to its infinite-dimensional mean-field limit. Our approach integrates Transformer architecture with mean-field limit theory and the Cucker–Smale flocking model. Contribution/Results: We rigorously derive an explicit upper bound on the approximation error—expressed in terms of model and system parameters—and prove that the framework is inherently permutation-equivariant and consistent with the underlying mean-field PDE. Empirical validation is conducted on flocking dynamics and mean-field neural network training tasks. Numerical results closely match the theoretical error bounds, confirming both practical efficacy and theoretical soundness.
This paper addresses the cooperative control of large-scale homogeneous agents in mean-field Markov decision processes (MFG-MDPs), focusing on optimizing a single-agent policy via reinforcement learning to minimize the aggregate social cost. For the mean-field linear-quadratic (MF-LQ) setting, we establish, for the first time, the global convergence of both exact and model-free policy gradient algorithms—providing the first theoretical convergence guarantee for mean-field reinforcement learning. Our approach integrates tools from linear-quadratic stochastic control, mean-field game modeling, and stochastic approximation theory, enabling convergence without prior knowledge of the environment dynamics. Numerical experiments confirm the predicted convergence rates and theoretical consistency. The key contribution is the first model-free policy gradient framework for distributed learning in large-scale agent systems with provable global convergence.
To address the inefficiency and instability of conventional forward-backward fixed-point iteration (FPI) in mean-field game (MFG) learning for large-scale multi-agent systems—caused by oscillatory behavior—this paper proposes a unified optimization framework that jointly treats policies and population distributions as co-optimizable control variables, enabling asynchronous joint updates. We introduce the first gradient-based MFG learning algorithm for continuous state-action spaces: population-aware linear function approximation (PA-LFA). Theoretically, we prove finite-time convergence to exact equilibria for contractive linear MFGs, asymptotic convergence to an equilibrium neighborhood under milder conditions, and derive approximation error bounds for nonlinear MFGs. Extensive evaluation across six benchmark tasks demonstrates the method’s effectiveness and robustness.
This paper investigates linear-quadratic (LQ) mean field games (MFGs) in infinite-dimensional Hilbert spaces, where each of the $N$ agents evolves according to an infinite-dimensional stochastic dynamics driven by a $Q$-Wiener process, features unbounded operators in its drift, and interacts via the empirical average of states. Methodologically, it establishes, for the first time, an infinite-dimensional Nash certainty equivalence principle, integrating operator semigroup theory, stochastic evolution equations, and mean-field asymptotic analysis to rigorously characterize the unique mean-field Nash equilibrium. Theoretical contributions include: (i) strong convergence of the empirical state average to the mean-field limit solution; (ii) construction of limiting optimal response strategies that constitute an $varepsilon$-Nash equilibrium for the $N$-player game; and (iii) establishment of a complete theoretical foundation for infinite-dimensional LQ-MFGs, providing a novel analytical framework for multi-agent systems governed by partial differential structures.
This paper studies an N-player stochastic game with irreversible investment and its corresponding mean-field game: each player controls a geometric Brownian motion state variable via nondecreasing singular controls to maximize long-run average power-type utility. Within an ergodic singular control framework, we pioneer the application of Lagrange multiplier methods to solve the mean-field control problem. We introduce the novel concept of “mean-field coarse correlated equilibrium” tailored to stationary settings, capturing strategic complementarities and interdependence among agents. Explicit constructions are provided for three types of equilibria—mean-field control, coarse correlated equilibrium, and Nash equilibrium—and we rigorously establish that, as (N o infty), both the coarse correlated and Nash equilibria converge to the mean-field solution. Numerical experiments further compare existence conditions and payoff differences across these equilibria.
This work addresses the computational intractability arising from complex agent interactions in large-scale multi-agent reinforcement learning by proposing a scalable learning framework grounded in mean-field control theory. By leveraging mean-field approximation to characterize population behavior, the approach constructs a representative agent model and integrates it with a Markov decision process subject to common noise, thereby establishing a rigorous theoretical link between finite-population systems and their mean-field limits. The study presents the first systematic unification of mean-field control and reinforcement learning, offering formal analyses of propagation of chaos and algorithmic convergence. It further incorporates dynamic programming, Q-learning, policy gradient, and DDPG methods within this framework. Empirical validation on both general and linear-quadratic models demonstrates the efficacy of the proposed algorithms in efficiently approximating solutions for large-scale stochastic multi-agent systems.
本文探讨了在大规模多智能体系统中,通过学习低维群体表示来解决高维控制问题,从而优化策略并提高奖励预测和纳什均衡质量。
本文探讨了控制理论、最优传输、概率推理、非平衡热力学和机器学习之间的联系,通过优化自由能类函数解决高维数据中的复杂结构学习问题。
This study addresses the limitation of the standard REINFORCE algorithm in capturing population distribution effects within mean-field control by proposing the Transport REINFORCE method. This approach integrates model-free policy gradients, optimal transport mapping theory, and Gaussian mixture projection techniques. By perturbing transformed distributions on the probability simplex or Gaussian mixture manifold, it effectively estimates the missing mean-field contributions and policy gradients, accommodating both finite and continuous state spaces. Theoretically, we establish the consistency of the proposed estimator and derive its error bounds. Empirically, experiments demonstrate that Transport REINFORCE significantly outperforms the standard REINFORCE baseline, validating its effectiveness for mean-field control problems.
This work addresses the challenges of modeling complex continuous dynamical systems—particularly strong nonlinearity, high-dimensional state spaces, and difficulties in uncertainty quantification—by proposing a novel framework that integrates Gaussian process ordinary differential equations with second-order reduced-order modeling. The approach learns latent-space dynamics through kernel-based autonomous ODEs and employs spherical projection to ensure numerical stability. Theoretically grounded with convergence guarantees, the method achieves substantially improved prediction accuracy and computational efficiency. Empirical evaluations on multiple benchmark systems demonstrate its superior performance over mainstream techniques such as extended dynamic mode decomposition, exhibiting greater robustness and practicality in terms of predictive accuracy, computational cost, and uncertainty quantification.