🤖 AI Summary
To address the challenges of modeling inter-product demand dependencies and weak coordination in retail dynamic pricing, this paper proposes Graph-Augmented Multi-Agent Proximal Policy Optimization (GAT-MAPPO). It constructs a product-relational graph and incorporates graph attention mechanisms to explicitly capture dynamic demand interactions among products, while leveraging MAPPO as the learning framework for end-to-end joint pricing policy optimization. Compared to independent learning and standard MAPPO baselines, GAT-MAPPO achieves significant improvements in a real-data-driven simulation: +12.3% higher overall profit, −38.7% lower price volatility, enhanced cross-product pricing fairness, and improved training stability. The core contribution lies in integrating structured domain priors—namely, product associations—into multi-agent reinforcement learning, thereby balancing scalability with policy consistency.
📝 Abstract
Dynamic pricing in retail requires policies that adapt to shifting demand while coordinating decisions across related products. We present a systematic empirical study of multi-agent reinforcement learning for retail price optimization, comparing a strong MAPPO baseline with a graph-attention-augmented variant (MAPPO+GAT) that leverages learned interactions among products. Using a simulated pricing environment derived from real transaction data, we evaluate profit, stability across random seeds, fairness across products, and training efficiency under a standardized evaluation protocol. The results indicate that MAPPO provides a robust and reproducible foundation for portfolio-level price control, and that MAPPO+GAT further enhances performance by sharing information over the product graph without inducing excessive price volatility. These results indicate that graph-integrated MARL provides a more scalable and stable solution than independent learners for dynamic retail pricing, offering practical advantages in multi-product decision-making.