Score
Design and implement algorithms and training procedures that optimize decentralized multi‑agent policies by training actors independently with serialized updates and by conditioning each agent’s policy on a belief or estimate of previous agents’ actions. Build methods to chain per‑agent updates into a coordinated joint‑gradient update and analyze the mechanisms and guarantees (e.g., convergence or improvement) that ensure the chained, decentralized updates produce joint‑policy improvement.
This work systematically investigates three canonical interaction paradigms in multi-agent reinforcement learning (MARL): federated cooperation, decentralized collaboration, and non-cooperative games—corresponding respectively to centralized coordination, transient interaction, and incentive conflicts. Method: We propose the first unified taxonomy integrating federated learning principles into MARL, formally characterizing the theoretical boundaries and modeling assumptions of each paradigm’s topology. Leveraging tools from Markov decision processes, distributed optimization, and game theory, we conduct rigorous theoretical analysis and empirical evaluation. Contribution/Results: Our study clarifies fundamental trade-offs across convergence guarantees, communication efficiency, and equilibrium stability among paradigms, and identifies shared bottlenecks—including heterogeneity, non-stationarity, and incentive incompatibility—in existing approaches. The framework provides a structured conceptual foundation for MARL interaction modeling and informs future research directions toward practical deployment.
This work addresses the challenge of efficiently computing joint policy gradients for global return maximization within the centralized training with decentralized execution (CTDE) framework. The authors propose a novel policy optimization method based on sequential joint decision-making, which provides the first exact decentralized decomposition of the joint policy gradient. By introducing an action belief mechanism to coordinate agent interactions and integrating sequential action commitments, decentralized critics, and individual score functions, each agent can perform independent updates while collectively implementing a complete joint gradient step. The approach does not rely on value factorization assumptions or converge to suboptimal equilibria. It achieves significant performance gains over strong baselines on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo benchmarks, with advantages amplifying as the number of agents scales up.
This work addresses decentralized multi-agent navigation in cluttered environments, proposing the first joint optimization framework for agent policies and reconfigurable environmental layouts (e.g., obstacle placements). Methodologically, it employs model-free policy gradient reinforcement learning and introduces a two-stage alternating optimization algorithm that concurrently updates distributed agent policies and environmental structure. Theoretical analysis establishes convergence to local minima of a time-varying non-convex optimization problem. A key finding is that the optimized environment autonomously forms implicit, motion-decoupled guidance structures—enhancing behavioral coordination without explicit communication or centralized control. Experiments across diverse dense scenarios demonstrate consistent superiority over baselines in navigation success rate, throughput efficiency, and collision rate, empirically validating that environmental configuration optimization delivers substantial gains for multi-agent collaborative navigation.
This paper addresses decentralized, asynchronous, communication-free, and model-free multi-agent reinforcement learning in infinite-horizon discounted Markov potential games. We propose a two-timescale asynchronous stochastic approximation framework that integrates local Q-function estimation with actor-critic–inspired policy updates, enabling decoupled learning using only individual reward observations. For the first time, we rigorously apply two-timescale analysis to establish almost-sure convergence of the learning dynamics to the set of Nash equilibria in this setting. Experiments demonstrate rapid convergence and robustness across standard potential game benchmarks. Our key contributions are: (1) a fully decentralized algorithm requiring no global information or coordination mechanisms; and (2) the first rigorous theoretical guarantee for the convergence of asynchronous Q-learning in decentralized Markov potential games. The analysis explicitly handles asynchrony, partial observability, and unknown environment dynamics while preserving equilibrium stability.
In decentralized execution for cooperative multi-agent reinforcement learning (MARL), mainstream decentralized policy gradient methods suffer from inherent suboptimality, preventing convergence to globally optimal policies. Method: We propose the Transformation-and-Distillation (TAD) framework, which equivalently reformulates a cooperative multi-agent MDP into a sequential single-agent MDP and employs policy distillation to recover decentralized execution. Contribution/Results: We theoretically prove that TAD guarantees learning of globally optimal policies in finite MDPs. Instantiating TAD with PPO, we develop TAD-PPO—incorporating MDP structural transformation, two-stage training, and value decomposition analysis. Empirical evaluation across diverse cooperative benchmarks demonstrates that TAD-PPO significantly outperforms state-of-the-art methods, achieving both theoretical global optimality guarantees and strong generalization capability.
Decentralized multi-agent reinforcement learning (MARL) suffers from non-stationarity due to asynchronous policy updates across agents, undermining convergence guarantees. Method: We propose a fully decentralized asynchronous Q-learning algorithm that eliminates the need for synchronization. For the first time, we establish a two-timescale stochastic approximation framework under constant learning rates, integrating Markov chain modeling with persistence-based convergence analysis—bypassing any reliance on synchronized policy updates. Contribution/Results: Under mild assumptions, we prove—with high probability—that agent policies converge to a Nash equilibrium. Our analytical framework generalizes beyond Q-learning to encompass multiple classes of regret-minimization algorithms. By removing synchronization requirements, the algorithm achieves significantly enhanced applicability and robustness in realistic asynchronous, decentralized environments—advancing both theoretical rigor and practical deployment of MARL.
This paper addresses the joint optimization of reward interdependence and policy coupling in networked multi-agent reinforcement learning (NMARL): each agent’s reward depends on its own and κₚ-hop neighbors’ state-action pairs, while its policy is parameterized jointly by its own and κₚ-hop neighbors’ parameters. To this end, we propose Distributed Coupled Policy Optimization (DCPO), a scalable decentralized algorithm that introduces neighbor-averaged Q-functions and coupled policy gradients, and employs a geometric two-step sampling scheme—obviating explicit Q-tables. DCPO integrates push-sum consensus for fully decentralized coordination. We establish theoretical convergence to first-order stationary points of the objective function under mild assumptions. Empirical evaluation on robotic path planning tasks demonstrates that DCPO significantly outperforms existing NMARL methods in terms of sample efficiency, scalability, and practical applicability, while preserving robustness to network topology changes.
In CTDE frameworks, the critic-decentralization mismatch (CDM) arises when centralized critics and decentralized actors are misaligned, causing agents to degrade in learning due to others’ suboptimal behaviors. Existing value decomposition methods face a trade-off: linear decompositions enable decentralized gradient computation but lack representational capacity; nonlinear decompositions offer stronger expressivity yet require centralized gradients, reintroducing CDM. This paper proposes Monotonic Nonlinear Critic Decomposition (MNCD) and Multi-Agent Cross-Entropy Optimization (MACE), enabling fully decentralized policy updates while significantly enhancing joint value function representation. Furthermore, we integrate an improved *k*-step return with Retrace off-policy correction to boost training stability and sample efficiency. Empirical evaluation demonstrates that our approach consistently outperforms state-of-the-art methods across multiple continuous- and discrete-action benchmark tasks.
This paper addresses the challenge of learning temporally coordinated multi-agent policies for multi-task settings under the centralized training with decentralized execution (CTDE) paradigm. To overcome the low sample efficiency and poor generalization of existing methods to diverse tasks, we propose ACC-MARL—a framework that models temporal tasks as finite-state automata to enable explicit task decomposition and agent coordination. It introduces a task-conditioned policy network and, within the CTDE framework, incorporates a value-function-driven online task assignment mechanism that dynamically optimizes role allocation during execution. Experiments demonstrate that ACC-MARL successfully emergently learns multi-step collaborative behaviors—such as cooperative door-opening and sequential unlocking—achieving significant improvements in task success rate, sample efficiency, and cross-task generalization.
This work addresses the lack of high-probability convergence guarantees in decentralized stochastic optimization, particularly under data heterogeneity and non-strongly convex settings where existing methods rely on overly restrictive assumptions. The paper proposes a gradient-tracking-based decentralized stochastic gradient descent algorithm (GT-DSGD) that, under mild sub-Gaussian noise conditions, establishes the first high-probability convergence guarantee for a bias-corrected decentralized method, thereby bridging the theoretical gap between high-probability and mean-square-error analyses. The analysis shows that GT-DSGD achieves optimal high-probability convergence rates of $O(\log(1/\delta)/\sqrt{nT})$ for non-convex objectives and $O(\log(1/\delta)/(nT))$ under the Polyak–Łojasiewicz condition, with experiments demonstrating its clear superiority over current state-of-the-art approaches.