decentralized policy synthesis

Designs and implements training algorithms and architectures that leverage centralized information, critics, or global signals during learning to produce per-agent decentralized policies that can execute using only local observations. This work includes building centralized critics and training objectives, defining how global information is accessed and distilled into local decision rules, and analyzing performance, coordination, and generalization of the resulting decentralized policies.

decentralizedpolicysynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A Communication-Efficient Decentralized Actor-Critic Algorithm

Oct 21, 2025
XR
Xiaoxing Ren
🏛️ Cornell University | KTH Royal Institute of Technology | Imperial College London

This paper addresses efficient multi-agent coordination in reinforcement learning under communication constraints. We propose a decentralized actor-critic framework that integrates local policy updates, multi-step local training, and sparse inter-agent information exchange, while employing multi-layer neural networks to approximate value functions—thereby reducing communication dependency. To our knowledge, this is the first work to establish finite-time convergence analysis under Markovian sampling, explicitly quantifying how neural network approximation error affects convergence accuracy. Theoretically, we prove that the algorithm achieves an ε-accurate stationary point with sample complexity O(ε⁻³) and communication complexity O(ε⁻¹τ⁻¹), where τ characterizes the mixing time of the underlying Markov process. Extensive experiments on cooperative control tasks validate the method’s superior empirical performance and strong alignment with the derived theoretical bounds.

Develops communication-efficient decentralized reinforcement learning for multi-agent systemsEstablishes finite-time convergence with neural network approximation analysisReduces communication burden while maintaining coordination through local policy updates

Fully-Decentralized MADDPG with Networked Agents

Mar 09, 2025
DB
Diego Bolliger
🏛️ ETH Zürich

This work addresses the computational bottleneck in multi-agent reinforcement learning (MARL) arising from centralized training in large-scale continuous action spaces. We propose a fully decentralized distributed training framework grounded in a networked communication topology. Departing from global centralized critics, our approach employs a distributed actor-critic architecture coupled with a surrogate policy mechanism, enabling cooperative policy optimization using only local neighbor communication. The framework natively supports cooperative, competitive, and mixed-task scenarios. Experiments demonstrate that our method matches MADDPG’s performance across diverse benchmark tasks while substantially reducing per-step training cost; this efficiency advantage scales favorably with increasing agent count. Overall, it establishes a new paradigm for scalable, low-overhead decentralized MARL.

Decentralized multi-agent reinforcement learning in continuous action spaces.Networked communication approach for decentralized MADDPG adaptation.Surrogate policies reduce computational cost in large-scale agent systems.

Is Centralized Training with Decentralized Execution Framework Centralized Enough for MARL?

May 27, 2023
YZ
Yihe Zhou
🏛️ Zhejiang University | China Electric Power Research Institute

Existing CTDE frameworks permit access to global state during training but suffer from insufficient exploitation of inter-agent cooperation cues and inefficient joint policy exploration due to enforced policy independence. To address this, we propose Centralized Advice with Decentralized Pruning (CADP), a novel paradigm that introduces an explicit cross-agent advice mechanism to facilitate efficient collaborative learning during training, while integrating differentiable smooth model pruning to eliminate redundant parameters and enhance policy consistency—without compromising fully decentralized execution. Evaluated on StarCraft II micromanagement and Google Research Football benchmarks, CADP consistently outperforms state-of-the-art CTDE methods, achieving significant improvements in joint policy exploration efficiency and cooperative generalization. Our approach provides a principled framework for enhancing multi-agent coordination under the CTDE paradigm.

CTDE framework limits global cooperative information sharingInefficient joint-policy exploration in current CTDE methodsNeed for decentralized execution with enhanced centralized training

Decentralized Learning Strategies for Estimation Error Minimization with Graph Neural Networks

Apr 04, 2024
XC
Xingran Chen
🏛️ University of Electronic Science and Technology of China | Duke University | University of Pennsylvania

This paper addresses distributed sampling and remote estimation of autoregressive Markov processes in multihop wireless networks, aiming to jointly minimize time-average estimation error and Age of Information (AoI). We tackle key challenges: statistically homogeneous agents, collision-prone shared channels, and the necessity of caching the latest sample. We establish, for the first time, that under a no-sensing policy, minimizing estimation error is equivalent to minimizing AoI. To this end, we propose a topology-transferable multi-agent reinforcement learning framework based on graph neural networks (GNNs), integrating permutation-equivariant architecture, recurrent state modeling, and centralized training with decentralized execution (CTDE). Experiments demonstrate that our approach significantly outperforms state-of-the-art methods across diverse network scales and dynamic environments. The learned policy exhibits strong cross-scale transferability, with performance gains increasing as node count grows. Moreover, the recurrent structure substantially improves robustness against non-stationary channel dynamics.

Develop decentralized scalable sampling and transmission policies.Minimize estimation error in multi-hop wireless networks.Optimize policies using graph neural networks for transferability.

Independent and Decentralized Learning in Markov Potential Games

May 29, 2022
CM
C. Maheshwari
🏛️ University of California, Berkeley

This paper addresses decentralized, asynchronous, communication-free, and model-free multi-agent reinforcement learning in infinite-horizon discounted Markov potential games. We propose a two-timescale asynchronous stochastic approximation framework that integrates local Q-function estimation with actor-critic–inspired policy updates, enabling decoupled learning using only individual reward observations. For the first time, we rigorously apply two-timescale analysis to establish almost-sure convergence of the learning dynamics to the set of Nash equilibria in this setting. Experiments demonstrate rapid convergence and robustness across standard potential game benchmarks. Our key contributions are: (1) a fully decentralized algorithm requiring no global information or coordination mechanisms; and (2) the first rigorous theoretical guarantee for the convergence of asynchronous Q-learning in decentralized Markov potential games. The analysis explicitly handles asynchrony, partial observability, and unknown environment dynamics while preserving equilibrium stability.

Discounted Markov Potential GamesMulti-Agent LearningStrategy Optimization

Latest Papers

What's happening recently
View more

This study investigates how agents in social networks achieve global coordination through local interactions and private signals, with a focus on how network structure shapes higher-order beliefs and equilibrium behavior. The authors develop a local-global game framework in which each agent makes decisions based on signals from neighbors within graph distance \( r \), and introduce the novel concept of “networked common learning” to characterize coordination efficiency. Integrating global game theory, social network analysis, and Bayesian inference, they employ graph-theoretic methods to analyze learning dynamics across different topologies. Their results show that structures such as two-dimensional grids support networked common learning and enable efficient coordination, whereas architectures with information bottlenecks—like linear chains—converge only to risk-dominant equilibria.

coordinationglobal gameshigher-order beliefs

This work challenges the conventional view that decentralized optimization is merely a suboptimal compromise necessitated by communication constraints or privacy requirements, and inherently inferior to centralized methods in convergence efficiency. The authors propose a novel decentralized optimization framework that, under a fair setting where each iteration incurs identical per-node computation and communication time, demonstrates for the first time that decentralized algorithms can achieve faster convergence in terms of iteration count than their centralized counterparts. Empirical evaluations on logistic regression and neural network training tasks consistently show accelerated learning, while simultaneously offering enhanced privacy preservation and improved system scalability. These results reframe decentralization not as a pragmatic fallback, but as a strategic choice that can yield tangible performance gains.

centralized learningconvergence accelerationdecentralized optimization

This work addresses the challenge of efficiently computing joint policy gradients for global return maximization within the centralized training with decentralized execution (CTDE) framework. The authors propose a novel policy optimization method based on sequential joint decision-making, which provides the first exact decentralized decomposition of the joint policy gradient. By introducing an action belief mechanism to coordinate agent interactions and integrating sequential action commitments, decentralized critics, and individual score functions, each agent can perform independent updates while collectively implementing a complete joint gradient step. The approach does not rely on value factorization assumptions or converge to suboptimal equilibria. It achieves significant performance gains over strong baselines on Multi-Robot Warehouse, SMACv2, and MA-MuJoCo benchmarks, with advantages amplifying as the number of agents scales up.

Centralized Training with Decentralized ExecutionCooperative TasksJoint Policy Optimization

Hot Scholars

HH

Hoda Heidari

Carnegie Mellon University
Responsible AIAI EthicsAI AccountabilityAlgorithmic Fairness
MD

Merouane Debbah

KU 6G Center, Khalifa University, Centralesupelec
6GLarge Language ModelsAIRandom Matrix Theory
GS

Guillaume Sartoretti

Assistant Professor, National University of Singapore (NUS), Mechanical Engineering Dpt
Multi-Agent SystemsRoboticsSwarm IntelligenceDistributed Control
AA

Alessandro Abate

Professor of Verification and Control, University of Oxford, UK
Formal VerificationControl TheoryStochastic Hybrid SystemsCyber-Physical Systems
YL

Yong Li

Institue of Software, Chinese Academy of Sciences
Automata theoryModel checking