Score
Designs and implements multi‑agent reinforcement learning algorithms and models that represent agents and their pairwise and higher‑order interactions as explicit relational or graph‑structured components, producing decentralized policies that reason about those relations. Builds and analyzes training and inference procedures that learn strategic interaction structure from observations and produce robust cooperative and competitive behaviors under partial observability and limited or no communication.
This work addresses the absence of a systematic taxonomy and unified framework for communication mechanisms in graph neural network (GNN)-driven multi-agent reinforcement learning. To bridge this gap, the paper proposes a general GNN-based communication pipeline and establishes the first structured survey and classification scheme, clearly delineating the core mechanisms and design principles underlying such approaches. By integrating insights from GNNs, multi-agent reinforcement learning, and communication modeling, this study enhances conceptual clarity and accessibility in the field, while also laying a theoretical foundation and offering methodological guidance for future research.
Existing large-scale multi-agent reinforcement learning (MARL) methods neglect structured inter-agent couplings, resulting in inefficient training and poor scalability. This paper proposes Partially Decentralized Training with Decentralized Execution (PD-TDE), a model-agnostic framework that models value dependencies among agents via Bayesian networks. It introduces the concept of value dependency sets into policy gradient theory for the first time and rigorously proves that the corresponding gradient estimator exhibits lower variance than centralized training. To enhance scalability, we design a truncated Bayesian network approximation algorithm that preserves convergence guarantees while significantly accelerating training. Evaluated on multi-warehouse resource scheduling and multi-zone thermal control tasks, PD-TDE improves training efficiency by 37%–52% over baselines, scales to hundreds of agents, and achieves a 2.1× speedup in convergence under high-coupling scenarios.
This work systematically investigates three canonical interaction paradigms in multi-agent reinforcement learning (MARL): federated cooperation, decentralized collaboration, and non-cooperative games—corresponding respectively to centralized coordination, transient interaction, and incentive conflicts. Method: We propose the first unified taxonomy integrating federated learning principles into MARL, formally characterizing the theoretical boundaries and modeling assumptions of each paradigm’s topology. Leveraging tools from Markov decision processes, distributed optimization, and game theory, we conduct rigorous theoretical analysis and empirical evaluation. Contribution/Results: Our study clarifies fundamental trade-offs across convergence guarantees, communication efficiency, and equilibrium stability among paradigms, and identifies shared bottlenecks—including heterogeneity, non-stationarity, and incentive incompatibility—in existing approaches. The framework provides a structured conceptual foundation for MARL interaction modeling and informs future research directions toward practical deployment.
To address the scalability bottleneck in multi-agent reinforcement learning (MARL) arising from the exponential growth of the joint action space with the number of agents, this paper proposes a centralized learning framework based on sequential abstraction. The core innovation is the introduction of a “supervisor” meta-agent that decouples the high-dimensional joint action into temporally ordered individual action sequences, thereby enabling structured dimensionality reduction of the joint action space. By decomposing the action space sequentially and parameterizing the joint policy with lightweight components, the method preserves global coordination while significantly reducing computational complexity. Experiments across multi-agent tasks of varying scales demonstrate that our approach outperforms existing centralized methods in training efficiency, convergence speed, and cooperative performance. Moreover, it exhibits strong generalization capability and practical deployability.
To address the four core challenges in multi-agent reinforcement learning (MARL)—non-stationarity, partial observability, large-scale scalability, and decentralized learning—this paper proposes a game-theoretic deep learning framework for cooperative learning. Methodologically, it unifies Nash equilibrium, evolutionary dynamics, correlated equilibrium, and adversarial dynamics into a single differentiable analytical paradigm, enabling gradient-based mapping from game-theoretic solutions to distributed policy updates. The framework integrates stochastic game modeling, projection-based policy-space gradient methods, population-level evolutionary differential equation approximations, and a decentralized actor-critic architecture. Empirically, on mixed cooperative-competitive benchmarks, it achieves a 42% improvement in convergence stability and a 37% gain in policy robustness. Moreover, it supports real-time Nash approximation for systems with up to one thousand agents. The approach bridges theoretical rigor—grounded in dynamic game theory—with engineering robustness, offering both formal guarantees and practical scalability.
This work addresses the computational intractability and observational uncertainty inherent in multi-agent reinforcement learning for partially observable stochastic games (POSGs). To mitigate these challenges, we propose a modeling paradigm grounded in inter-agent information sharing. By introducing abstractions of public information and imposing structured assumptions—such as joint observability and delayed sharing—we construct tractable approximations of POSGs. We establish, for the first time, a theoretical necessity of information sharing for overcoming the computational hardness of POSGs. Furthermore, we design a unified learning framework that jointly computes equilibrium policies and team-optimal solutions. Our algorithm achieves quasi-polynomial time and sample complexity. Notably, for cooperative POSGs, it yields the first rigorous statistical and computational complexity bounds for team-optimal solutions—achieving both statistical and computational quasi-efficiency.
Existing collaborative multi-agent reinforcement learning (MARL) frameworks lack flexible, scalable, and end-to-end trainable architectures, often relying on fixed topologies or centralized training paradigms. Method: We propose Reinforcement Networks (RN), the first MARL framework that unifies agent systems as arbitrary directed acyclic graphs (DAGs), enabling modular, hierarchical, and graph-structured coordination. RN introduces a DAG-driven agent organization, end-to-end gradient propagation across the graph, graph-aware policy optimization, a novel collaboration-aware credit assignment algorithm, and LevelEnv—an environment abstraction for reproducible evaluation. Contribution/Results: Experiments demonstrate that RN consistently outperforms state-of-the-art baselines across diverse cooperative MARL benchmarks, achieving simultaneous improvements in task performance, scalability, and structural expressiveness. RN establishes a new paradigm for structured, scalable MARL grounded in principled graph-based representation and learning.
This work addresses the inefficiencies in coordination and insufficient policy robustness in cooperative multi-agent reinforcement learning (MARL) that arise from relying solely on either local or global perspectives. To overcome this limitation, the paper proposes a Hierarchical Leader-Critic (HLC) architecture inspired by team organizational structures. HLC introduces, for the first time, a multi-level perspective mechanism into MARL, enabling synergistic learning of local and global information without explicit inter-agent communication by integrating high-level objectives with low-level execution. Coupled with a sequential training strategy, the proposed method significantly outperforms single-level baselines across multiple cooperative MARL benchmarks and demonstrates superior scalability and robustness as the number of agents and task complexity increase.
This work addresses the key challenges in multi-agent reinforcement learning—achieving efficient coordination, avoiding conflicts, and satisfying constraints—by proposing the Action Graph Policy (AGP). AGP enables decentralized decision-making by constructing an action dependency graph and a coordination context, allowing agents to reason over global action relationships. Theoretical analysis demonstrates that AGP’s joint policy representation is strictly more expressive than independent policies and surpasses existing centralized value decomposition approaches. Empirical results show that in partially observable tasks with anti-coordination penalties, AGP achieves success rates of 80–95%, substantially outperforming state-of-the-art MARL methods, which attain only 10–25%. Furthermore, AGP consistently leads across diverse multi-agent environments.
This paper addresses the challenge of learning temporally coordinated multi-agent policies for multi-task settings under the centralized training with decentralized execution (CTDE) paradigm. To overcome the low sample efficiency and poor generalization of existing methods to diverse tasks, we propose ACC-MARL—a framework that models temporal tasks as finite-state automata to enable explicit task decomposition and agent coordination. It introduces a task-conditioned policy network and, within the CTDE framework, incorporates a value-function-driven online task assignment mechanism that dynamically optimizes role allocation during execution. Experiments demonstrate that ACC-MARL successfully emergently learns multi-step collaborative behaviors—such as cooperative door-opening and sequential unlocking—achieving significant improvements in task success rate, sample efficiency, and cross-task generalization.
Learning implicit constraints from local generalized Nash equilibrium (GNE) demonstration data in multi-agent dynamic games remains challenging, especially under nonlinear dynamics and mixed convex/non-convex safety constraints. Method: We propose an inverse dynamic game-based framework for parametric constraint learning. Our approach encodes the KKT conditions of Nash equilibria as a mixed-integer linear program (MILP), enabling inner approximations of both convex and non-convex safe/unsafe sets—even under nonlinear system dynamics—for the first time. Contribution/Results: We theoretically prove that the learned constraint set is strictly contained within the true feasible set and characterize the learnability boundary under GNE demonstrations. Experiments on simulation and real-world robotic platforms demonstrate accurate recovery of diverse constraints and generation of robust, interaction-compliant motion trajectories. The method significantly improves planning safety and generalization in complex multi-agent scenarios.