decentralized deep rl

Designs and implements decentralized deep reinforcement learning agents and control policies that learn online from local observations and interactions; builds and analyzes systems for training and deploying deep RL policies that make real‑time control or resource‑allocation decisions without a central coordinator, addressing partial observability, limited communication, convergence, stability, and robustness.

decentralizeddeeprl

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenges of high sample complexity, absence of global information, and coordination difficulties in fully distributed multi-agent systems under networked discrete control. To this end, we propose a fully distributed multi-agent reinforcement learning algorithm incorporating safety constraints. The method introduces parallel neural network branches to process local observations, integrates a graph encoder with a state estimator to enhance local perception, and combines an Actor-Critic architecture with Control Barrier Functions to achieve built-in safety guarantees and policy constraints. Experimental results demonstrate that, compared to centralized policies, the proposed approach reduces average uncertainty by 26.3% while approximating the performance of computationally intensive attention-based algorithms with less than 1% error.

Decentralized ControlMulti-Agent Reinforcement LearningNetwork Control

This work addresses the lack of a scalable, decentralized multi-agent control framework for area coverage tasks that offers theoretical convergence guarantees. The authors propose a finite-state, decentralized policy-based control approach that decouples the coverage problem into two components: a reference-configuration-guided deep neural network design and a distributed control strategy based on agent policies. Central to their method are an anchor-follower triangular communication topology and the novel concept of Anyway Output Controllability (AOC). This framework yields a computationally efficient, time-invariant control policy capable of dynamically adapting to changes in the target region while theoretically ensuring convergence to the optimal coverage configuration, thereby achieving scalable and collaboratively efficient multi-agent coverage control.

coverage guaranteedecentralized controlfinite-state control

Decentralizing Multi-Agent Reinforcement Learning with Temporal Causal Information

Jun 09, 2025
JC
Jan Corazza
🏛️ University Alliance Ruhr | TU Dortmund University | Arizona State University

Decentralized Multi-Agent Reinforcement Learning (DMARL) faces fundamental challenges in ensuring policy compatibility, coping with communication constraints, and preserving privacy. Method: This paper proposes a temporally causal symbolic knowledge-guided distributed training framework. It pioneers the extension of formal policy compatibility verification to a symbolic knowledge space endowed with temporal-causal structure, integrating Symbolic AI, temporal logic modeling, multi-agent RL, and formal verification techniques. Contribution/Results: The framework enables scalable, decentralized training with theoretical guarantees—without requiring global state information or frequent inter-agent communication. It significantly improves policy compatibility and sample efficiency. Empirical evaluation on collaborative robot–drone tasks demonstrates accelerated convergence and higher task success rates.

Decentralized Multi-Agent RL faces policy compatibility challengesSymbolic knowledge aids privacy and communication constraintsTemporal causal information accelerates DMARL learning

Taxonomy and Trends in Reinforcement Learning for Robotics and Control Systems: A Structured Review

Oct 11, 2025
KT
Kumater Ter
🏛️ Air Force Institute of Technology | Ahmadu Bello University | FAMU-FSU College of Engineering

This paper addresses the persistent gap between theoretical advances in reinforcement learning (RL) and their practical deployment in robotics and control systems. To bridge this divide, we propose a structured taxonomy tailored to real-world robotic applications, grounded in the Markov decision process (MDP) framework and systematically incorporating mainstream deep RL algorithms—including DDPG, TD3, PPO, and SAC—across canonical domains such as motion control, dexterous manipulation, and multi-agent coordination. The taxonomy explicitly integrates training paradigms and deployment maturity metrics. Crucially, we identify recurring design patterns and evolutionary trends in high-dimensional continuous control tasks, thereby unifying theoretical insights with engineering constraints. Our framework advances reproducibility, transferability, and robustness in RL deployment on physical robots, offering both a methodological foundation and actionable guidelines for practitioners. (149 words)

Bridges theoretical RL advances with practical robotic implementationsCategorizes RL applications across locomotion and manipulation domainsReviews RL principles and algorithms for robotics and control systems

Synthesis of Hierarchical Controllers Based on Deep Reinforcement Learning Policies

Feb 21, 2024
FD
Florent Delgrange
🏛️ Vrije Universiteit Brussel | University of Antwerp | University of Haifa | Delft University of Technology | Aalborg University | Flanders Make

This paper addresses hierarchical control in a two-level environment: a high-level known graph (“map”) whose vertices represent unknown, dynamically evolving MDPs (“rooms”). We propose an end-to-end controller framework integrating deep reinforcement learning (DRL) with reactive synthesis. Methodologically, the low level trains reusable latent policies within each room using PAC learning theory—avoiding model distillation and improving robustness to sparse rewards; the high level employs reactive synthesis to generate a dynamic scheduler satisfying Linear Temporal Logic (LTL) specifications. Theoretically, we establish the first PAC performance guarantees for hierarchical policies and derive bounds on abstraction quality. Experimentally, on navigation tasks with dynamic obstacles, our framework significantly improves policy generalization across rooms and enhances reliability of high-level scheduling decisions.

Designing controllers for two-level structured environments with formal guaranteesEnsuring reusability and performance guarantees in high-level task planningTraining low-level policies without model distillation for scalability

Latest Papers

What's happening recently
View more

This study addresses the coordination challenge in decentralized inspection and maintenance planning for multi-component engineering systems by formulating the problem as a partially observable Markov decision process and employing multi-agent deep reinforcement learning (MADRL) approaches. The work encompasses a spectrum of training paradigms—from fully centralized to fully decentralized—including value decomposition and actor-critic architectures. A novel benchmark environment with tunable redundancy is introduced to systematically evaluate, for the first time, the coordination capabilities and policy optimality of various MADRL algorithms. Experimental results demonstrate that decentralized policies achieve near-optimal performance in low-redundancy series systems; however, coordination complexity increases significantly with higher redundancy. Despite this, all MADRL strategies consistently outperform an optimized heuristic baseline across tested scenarios.

coordinationdecentralizationinspection and maintenance planning

To address the poor interpretability of deep reinforcement learning (DRL) policies, this paper proposes a model-agnostic interpretability distillation framework. First, it adaptively partitions the state space using Voronoi diagrams; then, it distills locally linear policies on each partition. Unlike prior approaches, it imposes no structural assumptions, thus balancing interpretability and representational capacity while preserving policy transparency and achieving performance alignment—or even improvement. The key innovation lies in coupling geometric partitioning with knowledge distillation, endowing local linear models with both theoretical traceability and empirical effectiveness. Experiments on Gridworld and classical control benchmarks demonstrate that the distilled policies exhibit clear decision logic—e.g., piecewise-linear control laws—and achieve average performance gains of 1.2%–3.7% over the original DRL policies, significantly outperforming existing interpretable baselines.

Balancing model simplicity and accuracy in policy distillationCreating explainable RL policies from opaque deep neural networksPartitioning state space for specialized linear models using Voronoi

This work addresses the challenge of task offloading under stringent resource constraints in edge devices by proposing a decentralized deep reinforcement learning (DRL) agent that dynamically selects execution locations—local device, multi-access edge computing (MEC), or cloud—based on real-time conditions. The approach is the first to be deployed and evaluated on a real-world multi-device edge testbed integrated with live 5G communication, enabling a systematic assessment of the trade-offs between latency and energy consumption under local versus remote training paradigms. Experimental results demonstrate the feasibility of running DRL agents directly on end-user devices and quantitatively characterize the impact of different training deployment strategies on system performance, offering practical design guidelines for intelligent edge computing systems.

Decentralized Decision-MakingEdge ComputingOn-Device Learning

This work addresses the poor scalability and single-point-of-failure limitations of traditional centralized scheduling approaches in large-scale heterogeneous distributed systems, where dynamic workloads, resource heterogeneity, and competition for quality-of-service guarantees pose significant challenges. To overcome these issues, the authors propose DRL-MADRL, a fully decentralized multi-agent deep reinforcement learning framework that formulates task scheduling as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP). Leveraging a lightweight Actor-Critic architecture and relying solely on foundational libraries such as NumPy, the framework is efficiently deployable on resource-constrained edge devices. Experimental results on a 100-node system demonstrate that DRL-MADRL reduces average task completion time by 15.6% (30.8s vs. 36.5s), lowers energy consumption by 15.2% (745.2 vs. 878.3 kWh), and achieves an 82.3% SLA compliance rate (p < 0.001). The implementation is fully open-sourced to ensure reproducibility.

decentralized optimizationdistributed systemsheterogeneous resources

Existing multi-agent systems struggle to achieve flexible control through simple human instructions directed at only a subset of agents, and uncontrolled agents often lack the ability to autonomously infer and complement collaborative tasks. This work proposes a deep reinforcement learning–based multi-agent control framework that integrates policy following with an implicit task complementarity mechanism: by issuing commands solely to key agents, the remaining agents can adaptively infer their roles and autonomously complete the collaborative task. The approach overcomes the conventional requirement for global instructions and enables dynamic reconfiguration of collaboration structures to align with human intent. Experimental results demonstrate that the proposed method outperforms existing baselines in both collaborative flexibility and task performance.

controllabilityhuman-in-the-loop controlimplicit coordination

Hot Scholars

LD

Linkang Du

Xi'an Jiaotong University
Trustworthy Machine LearningDifferential Privacy
CZ

Chunyi Zhou

Zhejiang University
Cyberspace SecurityMachine Learning PrivacyFederated Learning
JC

Jiahao Chen

Zhejiang University
AI SecurityTrustworthy AIGenAI SecurityGenAI Privacy
YD

Yang Dai

Shenzhen Institutes of Advanced Technology,Chinese Academy of Sciences
perovskitesmemristor
SJ

Shouling Ji

Professor, Zhejiang University & Georgia Institute of Technology
Data-driven SecurityAI SecuritySoftware ScurityPrivacy