Fair Policy Optimization in Major-Minor Weakly Coupled Markov Decision Processes

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of fair resource allocation in weakly coupled Markov decision processes with primary and secondary agents, aiming to replace conventional utilitarian objectives with monotonically concave fairness functions. Theoretically, we prove that under symmetry conditions, fairness optimization can be reduced to a specific utilitarian objective. Methodologically, we propose a deep reinforcement learning algorithm based on counting proportions, integrated with a prioritized sampler to achieve efficient solutions. Experimental evaluations on machine replacement and taxi dispatching tasks demonstrate that the proposed approach exhibits both favorable scalability and strong fairness performance.
📝 Abstract
We consider fair resource allocation in sequential decision-making environments modeled as major-minor weakly coupled Markov decision processes (M2WCMDP). In this framework, resource constraints couple the action spaces of a major sub-Markov decision process (sub-MDP) and a population of minor sub-MDPs that would otherwise operate independently. Instead of using the traditional utilitarian (total-sum) objective, we optimize a general class of monotone, concave, permutation-invariant, normalized fairness functions. With homogeneous minor sub-MDPs, we prove that the problem under symmetry reduces to optimizing the platform-plus-mean-participant utilitarian objective over the class of \textit{permutation-invariant} policies, which allows us to exploit efficient algorithms that optimize the utilitarian-based objective to solve this fairness-aware problem. For more general settings, we introduce a count-proportion-based deep reinforcement learning approach with a priority-based sampler that generates feasible count actions. The generality of our framework means that the proposed algorithms and theoretical guarantees transfer to any domain with a symmetric M2WCMDP structure. We consider two applications: the machine replacement problem and the joint control of pricing and taxi relocation problem on a New York City-calibrated dataset. We validate our theoretical findings with comprehensive experiments, confirming the effectiveness of our proposed method in achieving strong fairness-aware performance while remaining scalable.
Problem

Research questions and friction points this paper is trying to address.

Fair resource allocation
Weakly coupled Markov decision processes
Sequential decision-making
Fairness optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fair Policy Optimization
Weakly Coupled MDPs
Deep Reinforcement Learning
Permutation-Invariant Policies
Count-Proportion-Based Sampling
🔎 Similar Papers
No similar papers found.