agent-chained policy optimization

Design and implement algorithms and training procedures that optimize decentralized multi‑agent policies by training actors independently with serialized updates and by conditioning each agent’s policy on a belief or estimate of previous agents’ actions. Build methods to chain per‑agent updates into a coordinated joint‑gradient update and analyze the mechanisms and guarantees (e.g., convergence or improvement) that ensure the chained, decentralized updates produce joint‑policy improvement.

agent-chainedpolicyoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of efficiently computing joint policy gradients for global return maximization within the centralized training with decentralized execution (CTDE) framework. The authors propose a novel policy optimization method based on sequential joint decision-making, which provides the first exact decentralized decomposition of the joint policy gradient. By introducing an action belief mechanism to coordinate agent interactions and integrating sequential action commitments, decentralized critics, and individual score functions, each agent can perform independent updates while collectively implementing a complete joint gradient step. The approach does not rely on value factorization assumptions or converge to suboptimal equilibria. It achieves significant performance gains over strong baselines on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo benchmarks, with advantages amplifying as the number of agents scales up.

Centralized Training with Decentralized ExecutionCooperative TasksJoint Policy Optimization

This work addresses decentralized multi-agent navigation in cluttered environments, proposing the first joint optimization framework for agent policies and reconfigurable environmental layouts (e.g., obstacle placements). Methodologically, it employs model-free policy gradient reinforcement learning and introduces a two-stage alternating optimization algorithm that concurrently updates distributed agent policies and environmental structure. Theoretical analysis establishes convergence to local minima of a time-varying non-convex optimization problem. A key finding is that the optimized environment autonomously forms implicit, motion-decoupled guidance structures—enhancing behavioral coordination without explicit communication or centralized control. Experiments across diverse dense scenarios demonstrate consistent superiority over baselines in navigation success rate, throughput efficiency, and collision rate, empirically validating that environmental configuration optimization delivers substantial gains for multi-agent collaborative navigation.

Co-optimize agent policies and reconfigurable environments for navigationDecentralized multi-agent navigation in cluttered reconfigurable spacesModel-free learning to improve agent-environment performance synergy

Independent and Decentralized Learning in Markov Potential Games

May 29, 2022
CM
C. Maheshwari
🏛️ University of California, Berkeley

This paper addresses decentralized, asynchronous, communication-free, and model-free multi-agent reinforcement learning in infinite-horizon discounted Markov potential games. We propose a two-timescale asynchronous stochastic approximation framework that integrates local Q-function estimation with actor-critic–inspired policy updates, enabling decoupled learning using only individual reward observations. For the first time, we rigorously apply two-timescale analysis to establish almost-sure convergence of the learning dynamics to the set of Nash equilibria in this setting. Experiments demonstrate rapid convergence and robustness across standard potential game benchmarks. Our key contributions are: (1) a fully decentralized algorithm requiring no global information or coordination mechanisms; and (2) the first rigorous theoretical guarantee for the convergence of asynchronous Q-learning in decentralized Markov potential games. The analysis explicitly handles asynchrony, partial observability, and unknown environment dynamics while preserving equilibrium stability.

Discounted Markov Potential GamesMulti-Agent LearningStrategy Optimization

Towards Global Optimality in Cooperative MARL with the Transformation And Distillation Framework

Jul 12, 2022
JY
Jianing Ye
🏛️ Washington University in St. Louis | Tsinghua University

In decentralized execution for cooperative multi-agent reinforcement learning (MARL), mainstream decentralized policy gradient methods suffer from inherent suboptimality, preventing convergence to globally optimal policies. Method: We propose the Transformation-and-Distillation (TAD) framework, which equivalently reformulates a cooperative multi-agent MDP into a sequential single-agent MDP and employs policy distillation to recover decentralized execution. Contribution/Results: We theoretically prove that TAD guarantees learning of globally optimal policies in finite MDPs. Instantiating TAD with PPO, we develop TAD-PPO—incorporating MDP structural transformation, two-stage training, and value decomposition analysis. Empirical evaluation across diverse cooperative benchmarks demonstrates that TAD-PPO significantly outperforms state-of-the-art methods, achieving both theoretical global optimality guarantees and strong generalization capability.

Addresses suboptimality in decentralized MARL algorithmsEnsures decentralized execution with theoretical performance guaranteesProposes a framework for global optimality in cooperative tasks

Asynchronous Decentralized Q-Learning: Two Timescale Analysis By Persistence

Aug 07, 2023
BY
Bora Yongacoglu
🏛️ Queen's University | University of Hawaii at Manoa

Decentralized multi-agent reinforcement learning (MARL) suffers from non-stationarity due to asynchronous policy updates across agents, undermining convergence guarantees. Method: We propose a fully decentralized asynchronous Q-learning algorithm that eliminates the need for synchronization. For the first time, we establish a two-timescale stochastic approximation framework under constant learning rates, integrating Markov chain modeling with persistence-based convergence analysis—bypassing any reliance on synchronized policy updates. Contribution/Results: Under mild assumptions, we prove—with high probability—that agent policies converge to a Nash equilibrium. Our analytical framework generalizes beyond Q-learning to encompass multiple classes of regret-minimization algorithms. By removing synchronization requirements, the algorithm achieves significantly enhanced applicability and robustness in realistic asynchronous, decentralized environments—advancing both theoretical rigor and practical deployment of MARL.

Addresses non-stationarity in multi-agent reinforcement learning.Proposes unsynchronized decentralized Q-learning for stochastic games.Relaxes synchronization assumptions using constant learning rates.

Latest Papers

What's happening recently
View more

Distributed scalable coupled policy algorithm for networked multi-agent reinforcement learning

Dec 05, 2025
PD
Pengcheng Dai
🏛️ Singapore University of Technology and Design | University of California, Riverside | Southeast University | Purple Mountain Laboratories

This paper addresses the joint optimization of reward interdependence and policy coupling in networked multi-agent reinforcement learning (NMARL): each agent’s reward depends on its own and κₚ-hop neighbors’ state-action pairs, while its policy is parameterized jointly by its own and κₚ-hop neighbors’ parameters. To this end, we propose Distributed Coupled Policy Optimization (DCPO), a scalable decentralized algorithm that introduces neighbor-averaged Q-functions and coupled policy gradients, and employs a geometric two-step sampling scheme—obviating explicit Q-tables. DCPO integrates push-sum consensus for fully decentralized coordination. We establish theoretical convergence to first-order stationary points of the objective function under mild assumptions. Empirical evaluation on robotic path planning tasks demonstrates that DCPO significantly outperforms existing NMARL methods in terms of sample efficiency, scalability, and practical applicability, while preserving robustness to network topology changes.

Addresses collaborative policy optimization using local neighbor interactions and scalable gradient estimation.Develops a distributed algorithm for multi-agent reinforcement learning with interdependent rewards and policies.Ensures convergence to optimal policies in networked agents without full global information.

In CTDE frameworks, the critic-decentralization mismatch (CDM) arises when centralized critics and decentralized actors are misaligned, causing agents to degrade in learning due to others’ suboptimal behaviors. Existing value decomposition methods face a trade-off: linear decompositions enable decentralized gradient computation but lack representational capacity; nonlinear decompositions offer stronger expressivity yet require centralized gradients, reintroducing CDM. This paper proposes Monotonic Nonlinear Critic Decomposition (MNCD) and Multi-Agent Cross-Entropy Optimization (MACE), enabling fully decentralized policy updates while significantly enhancing joint value function representation. Furthermore, we integrate an improved *k*-step return with Retrace off-policy correction to boost training stability and sample efficiency. Empirical evaluation demonstrates that our approach consistently outperforms state-of-the-art methods across multiple continuous- and discrete-action benchmark tasks.

Addresses centralized-decentralized mismatch in multi-agent reinforcement learningImproves sample efficiency with modified returns and off-policy learningOvercomes trade-off between gradient flexibility and representation expressiveness

Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning

Nov 04, 2025
BY
Beyazit Yalcinkaya
🏛️ University of California, Berkeley | Nissan Advanced Technology Center

This paper addresses the challenge of learning temporally coordinated multi-agent policies for multi-task settings under the centralized training with decentralized execution (CTDE) paradigm. To overcome the low sample efficiency and poor generalization of existing methods to diverse tasks, we propose ACC-MARL—a framework that models temporal tasks as finite-state automata to enable explicit task decomposition and agent coordination. It introduces a task-conditioned policy network and, within the CTDE framework, incorporates a value-function-driven online task assignment mechanism that dynamically optimizes role allocation during execution. Experiments demonstrate that ACC-MARL successfully emergently learns multi-step collaborative behaviors—such as cooperative door-opening and sequential unlocking—achieving significant improvements in task success rate, sample efficiency, and cross-task generalization.

Enabling emergent task-aware coordination under decentralized executionLearning multi-task multi-agent policies for temporal cooperative objectivesOvercoming sample inefficiency in automata-based task decomposition methods

This work addresses the lack of high-probability convergence guarantees in decentralized stochastic optimization, particularly under data heterogeneity and non-strongly convex settings where existing methods rely on overly restrictive assumptions. The paper proposes a gradient-tracking-based decentralized stochastic gradient descent algorithm (GT-DSGD) that, under mild sub-Gaussian noise conditions, establishes the first high-probability convergence guarantee for a bias-corrected decentralized method, thereby bridging the theoretical gap between high-probability and mean-square-error analyses. The analysis shows that GT-DSGD achieves optimal high-probability convergence rates of $O(\log(1/\delta)/\sqrt{nT})$ for non-convex objectives and $O(\log(1/\delta)/(nT))$ under the Polyak–Łojasiewicz condition, with experiments demonstrating its clear superiority over current state-of-the-art approaches.

bias-correctiondecentralized stochastic optimizationgradient tracking

Hot Scholars

MB

Murat Bronz

Assistant Professor of Applied Aerodynamics and UAV Systems, ENAC, Toulouse, France
Aircraft DesignAerial RoboticsEmbodied IntelligenceMDO
OA

Omran Ayoub

Lecturer-Researcher at SUPSI
Communication NetworksNetwork OptimizationMachine LearningExplainable AI
SW

Shijie Wang

PhD Candidate, The Hong Kong Polytechnic University
Graph Neural NetworksRecommender SystemsLarge Language ModelsTrustworthy AI
CL

Chengyi Liu

PhD of The Hong Kong Polytechnic University
Recommender SystemDiffusion ModelGNN
DB

Dominik Baumann

Aalto University, Espoo, Finland
Control TheoryRoboticsMachine LearningMulti-agent Systems