Policy Synthesis for Finite Populations of MDP Agents under Aggregate Reach-Avoid Chance Constraints

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses policy synthesis for finite-population Markov decision processes under aggregate chance constraints, where mean-field methods often violate constraints by neglecting stochastic fluctuations. The proposed approach propagates the second-order moments of the empirical density via discrete-time Lyapunov recursions and applies Cantelli’s inequality to reformulate chance constraints as deterministic conditions. The resulting optimal policies are computed by integrating sequential convex approximation with gradient-based methods. A key contribution is the development of moment-based surrogate constraints accompanied by rigorous finite-N safety certificates. Experiments on grid-world navigation and electric vehicle charging scenarios demonstrate that this method outperforms standard deterministic linear programming baselines, achieving high operational efficiency while strictly satisfying probabilistic safety constraints.
📝 Abstract
Consider a finite population of agents with decoupled Markov transition dynamics and empirical-density feedback, subject to the following constraints: with probability at least $1-δ_r$, at least a fraction $α_r$ of agents must reach a target region at some time $t^*$, while, at each time up to $t^*$, the unsafe population fraction must remain below $β_u$ with probability at least $1-δ_u$. However, standard mean-field methods enforce these constraints only in expectation, which fails to account for stochastic fluctuations at finite fleet size $N$. To address this control problem, we propagate the second-order moment (variance) of the empirical density alongside the mean-field trajectory via a discrete-time Lyapunov recursion, and apply the Cantelli inequality to convert chance constraints into tractable deterministic conditions on the moments of the empirical density. We then incorporate these moment-based surrogate constraints into a gradient-based sequential convex approximation procedure for density-feedback policy synthesis. We further introduce additional moment-error bounds to construct a rigorous finite-$N$ certificate. The method is evaluated on a gridworld environment and a power-system EV-charging aggregation problem and compared with a standard deterministic population-level LP baseline.
Problem

Research questions and friction points this paper is trying to address.

Policy Synthesis
Finite Population MDP
Chance Constraints
Mean-Field Methods
Reach-Avoid
Innovation

Methods, ideas, or system contributions that make the work stand out.

Finite population MDP
Chance constraints
Second-order moment propagation
Cantelli inequality
Sequential convex approximation