Score
Designs, trains, and evaluates multi‑agent market‑making systems that learn interacting pricing and quoting policies for multiple market‑maker agents; implements centralized training with decentralized execution (CTDE) and multi‑agent proximal policy optimization (MAPPO) to produce heterogeneous agents, handle stochastic training and complex market dynamics, and analyze profitability and robustness under variations in informedness.
Existing CTDE methods fail to fully exploit the advantages of centralized training and lack theoretical guarantees. This paper proposes MAGPO, a novel CTDE framework that enables efficient cooperative exploration and decentralized execution via centralized guidance and decentralized policy alignment. Its core contributions are: (1) an autoregressive joint policy modeling mechanism that supports scalable multi-agent cooperative exploration; and (2) a policy alignment constraint coupled with monotonic policy gradient optimization, establishing—for the first time—theoretical guarantees of monotonic policy improvement in CTDE. Evaluated across six environments and 43 tasks, MAGPO consistently outperforms mainstream CTDE baselines and achieves performance on par with or superior to fully centralized approaches, demonstrating its effectiveness, generalizability, and practicality.
Cooperative multi-agent reinforcement learning (Cooperative MARL) suffers from conceptual ambiguity regarding fundamental paradigms—particularly the distinctions and applicability boundaries among centralized training with centralized execution (CTE), centralized training with decentralized execution (CTDE), and fully decentralized training and execution (DTE)—under the common setting of global reward sharing. Method: This work establishes a unified analytical framework to systematically characterize the design principles, intrinsic relationships, and evolutionary trajectories of major approaches, including value-decomposition methods (e.g., VDN, QMIX, QPLEX) and centralized-critic methods (e.g., MADDPG, COMA, MAPPO). Contribution: The analysis rigorously clarifies long-standing conceptual confusions in Cooperative MARL, yielding a structured cognitive map that supports principled algorithm selection, fair method comparison, and informed investigation of open challenges. The framework serves both as a pedagogical tool for teaching and a foundational reference for research advancement.
Existing CTDE frameworks permit access to global state during training but suffer from insufficient exploitation of inter-agent cooperation cues and inefficient joint policy exploration due to enforced policy independence. To address this, we propose Centralized Advice with Decentralized Pruning (CADP), a novel paradigm that introduces an explicit cross-agent advice mechanism to facilitate efficient collaborative learning during training, while integrating differentiable smooth model pruning to eliminate redundant parameters and enhance policy consistency—without compromising fully decentralized execution. Evaluated on StarCraft II micromanagement and Google Research Football benchmarks, CADP consistently outperforms state-of-the-art CTDE methods, achieving significant improvements in joint policy exploration efficiency and cooperative generalization. Our approach provides a principled framework for enhancing multi-agent coordination under the CTDE paradigm.
To address the challenges of modeling inter-product demand dependencies and weak coordination in retail dynamic pricing, this paper proposes Graph-Augmented Multi-Agent Proximal Policy Optimization (GAT-MAPPO). It constructs a product-relational graph and incorporates graph attention mechanisms to explicitly capture dynamic demand interactions among products, while leveraging MAPPO as the learning framework for end-to-end joint pricing policy optimization. Compared to independent learning and standard MAPPO baselines, GAT-MAPPO achieves significant improvements in a real-data-driven simulation: +12.3% higher overall profit, −38.7% lower price volatility, enhanced cross-product pricing fairness, and improved training stability. The core contribution lies in integrating structured domain priors—namely, product associations—into multi-agent reinforcement learning, thereby balancing scalability with policy consistency.
This study investigates how the degree of market information availability affects the profitability of heterogeneous market makers. The authors develop a multi-agent market model that integrates endogenous price formation, self-exciting order flow modeled via a Hawkes process, and market makers’ heterogeneous information sets and risk preferences. Methodologically, they provide the first stability guarantee for state-dependent Hawkes-based order execution processes over a finite time horizon and employ the MAPPO algorithm under a centralized training with decentralized execution (CTDE) framework to derive optimal market-making strategies. The findings reveal that insufficient information exposes market makers to substantial adverse selection risk due to informed trading; however, as market informativeness increases, their aggregate profitability rises significantly, effectively offsetting this risk and underscoring the critical role of the information environment in shaping market microstructure efficiency.
This work addresses the high variance in advantage estimation caused by non-stationary teammate policies in multi-agent reinforcement learning, which undermines the effectiveness of ratio-based trust-region methods such as MAPPO and MASPO. To mitigate this issue, the authors propose MARS, a novel policy optimization objective that replaces conventional additive ratio clipping or soft penalty mechanisms with a multiplicative symmetric geometric barrier within the centralized training with decentralized execution (CTDE) framework. This design imposes unbounded penalties on probability ratios approaching zero while preserving informative gradients, thereby preventing policy collapse and vanishing gradients. Empirical evaluation across 47 tasks spanning eight benchmark environments demonstrates that MARS consistently matches or outperforms existing methods, and ablation studies confirm the critical role of the symmetric geometric barrier in its performance gains.
This work addresses the challenge of efficiently computing joint policy gradients for global return maximization within the centralized training with decentralized execution (CTDE) framework. The authors propose a novel policy optimization method based on sequential joint decision-making, which provides the first exact decentralized decomposition of the joint policy gradient. By introducing an action belief mechanism to coordinate agent interactions and integrating sequential action commitments, decentralized critics, and individual score functions, each agent can perform independent updates while collectively implementing a complete joint gradient step. The approach does not rely on value factorization assumptions or converge to suboptimal equilibria. It achieves significant performance gains over strong baselines on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo benchmarks, with advantages amplifying as the number of agents scales up.
In spatial public goods games, strong payoff coupling, environmental non-stationarity, and population-level strategy dependencies pose significant challenges for existing approaches—such as evolutionary updating or independent reinforcement learning—which fail to capture inter-agent strategic interdependence. To address this, this paper introduces multi-agent proximal policy optimization (MAPPO) to the domain for the first time, proposing the MAPPO-LCR framework. Its key contributions are: (1) a centralized critic that explicitly models global policy coupling; (2) a local cooperation reward (LCR) mechanism, where reward signals are derived from neighborhood cooperation density to guide policy updates; and (3) unified decoupled execution with joint value estimation, preserving the original game structure. Experiments across diverse enhancement factors demonstrate that MAPPO-LCR consistently emergently fosters cooperation, substantially outperforming independent PPO in cooperation rate, convergence speed, and robustness.
This study addresses the systemic risks posed by large language models acting as autonomous economic agents in multi-agent markets, including market instability from algorithmic feedback loops and trust erosion due to Sybil attacks. The authors introduce the concept of “economic alignment” and propose a four-dimensional quantitative metric, EAS, to evaluate it. Through the Agent Bazaar multi-agent simulation framework, they demonstrate that economic alignment is orthogonal to general-purpose capabilities. By integrating REINFORCE++ reinforcement learning, adaptive curriculum learning, and alignment mechanisms—Stabilizing Firms and Skeptical Guardians—they train a 9B-parameter specialized model that significantly outperforms current state-of-the-art and open-source models in stability, integrity, social welfare, and profitability. These results validate that targeted reinforcement learning can effectively enhance economic alignment.
This work proposes the first falsifiable and reproducible synthetic experimental framework for systematically comparing the coordination efficacy of centralized planning and polycentric market mechanisms within a unified simulated economic environment. The framework integrates input-output networks, heterogeneous firms, capacity constraints, and endogenous pricing, leveraging agent-based modeling, adversarial stress testing, and structural shock analysis. Experimental results demonstrate that computational planners consistently achieve lower welfare losses across training, holdout, and adversarial scenarios, thereby validating the framework’s effectiveness. This approach establishes a methodological prototype for empirical calibration and mechanism design research in comparative economic systems.