Randomized Transport Maps for Model-Free Policy-Gradient Mean-Field Control

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of the standard REINFORCE algorithm in capturing population distribution effects within mean-field control by proposing the Transport REINFORCE method. This approach integrates model-free policy gradients, optimal transport mapping theory, and Gaussian mixture projection techniques. By perturbing transformed distributions on the probability simplex or Gaussian mixture manifold, it effectively estimates the missing mean-field contributions and policy gradients, accommodating both finite and continuous state spaces. Theoretically, we establish the consistency of the proposed estimator and derive its error bounds. Empirically, experiments demonstrate that Transport REINFORCE significantly outperforms the standard REINFORCE baseline, validating its effectiveness for mean-field control problems.
📝 Abstract
We develop a model-free policy gradient method for discrete-time mean-field control (MFC). In MFC, the policy affects the objective both through the controlled dynamics and through the population distribution. Standard REINFORCE estimators capture the first effect but not the second. We introduce Transport REINFORCE, a transport map-based approach that perturbs a suitable transformation of the population distribution to estimate this missing mean-field contribution. The method applies to both finite and continuous state spaces. In finite state spaces, we perturb the population distribution directly on the probability simplex through a convex combination of the current population weights and random weights. In continuous state spaces, we project the population distribution onto the manifold of Gaussian mixtures, and then randomize it via a transport map that ensures the perturbed law remains within this manifold. We prove consistency of the perturbed objective and gradient as the perturbation vanishes, and derive bias and mean-square error bounds for the resulting sample-based gradient estimator. Numerical experiments on several MFC benchmarks show that Transport REINFORCE improves over standard REINFORCE.
Problem

Research questions and friction points this paper is trying to address.

mean-field control
policy gradient
model-free reinforcement learning
REINFORCE estimator
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mean-Field Control
Policy Gradient
Transport Maps
Model-Free Reinforcement Learning
REINFORCE
🔎 Similar Papers
No similar papers found.
A
Adonis Jamal
ENS Paris-Saclay
S
Samy Mekkaoui
École Polytechnique
Y
Yadh Hafsi
École Polytechnique
H
Huyên Pham
École Polytechnique