Score
Designs and implements decentralized deep reinforcement learning agents and control policies that learn online from local observations and interactions; builds and analyzes systems for training and deploying deep RL policies that make real‑time control or resource‑allocation decisions without a central coordinator, addressing partial observability, limited communication, convergence, stability, and robustness.
This work systematically investigates three canonical interaction paradigms in multi-agent reinforcement learning (MARL): federated cooperation, decentralized collaboration, and non-cooperative games—corresponding respectively to centralized coordination, transient interaction, and incentive conflicts. Method: We propose the first unified taxonomy integrating federated learning principles into MARL, formally characterizing the theoretical boundaries and modeling assumptions of each paradigm’s topology. Leveraging tools from Markov decision processes, distributed optimization, and game theory, we conduct rigorous theoretical analysis and empirical evaluation. Contribution/Results: Our study clarifies fundamental trade-offs across convergence guarantees, communication efficiency, and equilibrium stability among paradigms, and identifies shared bottlenecks—including heterogeneity, non-stationarity, and incentive incompatibility—in existing approaches. The framework provides a structured conceptual foundation for MARL interaction modeling and informs future research directions toward practical deployment.
This study addresses the challenges of high sample complexity, absence of global information, and coordination difficulties in fully distributed multi-agent systems under networked discrete control. To this end, we propose a fully distributed multi-agent reinforcement learning algorithm incorporating safety constraints. The method introduces parallel neural network branches to process local observations, integrates a graph encoder with a state estimator to enhance local perception, and combines an Actor-Critic architecture with Control Barrier Functions to achieve built-in safety guarantees and policy constraints. Experimental results demonstrate that, compared to centralized policies, the proposed approach reduces average uncertainty by 26.3% while approximating the performance of computationally intensive attention-based algorithms with less than 1% error.
This work addresses the lack of a scalable, decentralized multi-agent control framework for area coverage tasks that offers theoretical convergence guarantees. The authors propose a finite-state, decentralized policy-based control approach that decouples the coverage problem into two components: a reference-configuration-guided deep neural network design and a distributed control strategy based on agent policies. Central to their method are an anchor-follower triangular communication topology and the novel concept of Anyway Output Controllability (AOC). This framework yields a computationally efficient, time-invariant control policy capable of dynamically adapting to changes in the target region while theoretically ensuring convergence to the optimal coverage configuration, thereby achieving scalable and collaboratively efficient multi-agent coverage control.
Decentralized Multi-Agent Reinforcement Learning (DMARL) faces fundamental challenges in ensuring policy compatibility, coping with communication constraints, and preserving privacy. Method: This paper proposes a temporally causal symbolic knowledge-guided distributed training framework. It pioneers the extension of formal policy compatibility verification to a symbolic knowledge space endowed with temporal-causal structure, integrating Symbolic AI, temporal logic modeling, multi-agent RL, and formal verification techniques. Contribution/Results: The framework enables scalable, decentralized training with theoretical guarantees—without requiring global state information or frequent inter-agent communication. It significantly improves policy compatibility and sample efficiency. Empirical evaluation on collaborative robot–drone tasks demonstrates accelerated convergence and higher task success rates.
This paper addresses the persistent gap between theoretical advances in reinforcement learning (RL) and their practical deployment in robotics and control systems. To bridge this divide, we propose a structured taxonomy tailored to real-world robotic applications, grounded in the Markov decision process (MDP) framework and systematically incorporating mainstream deep RL algorithms—including DDPG, TD3, PPO, and SAC—across canonical domains such as motion control, dexterous manipulation, and multi-agent coordination. The taxonomy explicitly integrates training paradigms and deployment maturity metrics. Crucially, we identify recurring design patterns and evolutionary trends in high-dimensional continuous control tasks, thereby unifying theoretical insights with engineering constraints. Our framework advances reproducibility, transferability, and robustness in RL deployment on physical robots, offering both a methodological foundation and actionable guidelines for practitioners. (149 words)
This paper addresses hierarchical control in a two-level environment: a high-level known graph (“map”) whose vertices represent unknown, dynamically evolving MDPs (“rooms”). We propose an end-to-end controller framework integrating deep reinforcement learning (DRL) with reactive synthesis. Methodologically, the low level trains reusable latent policies within each room using PAC learning theory—avoiding model distillation and improving robustness to sparse rewards; the high level employs reactive synthesis to generate a dynamic scheduler satisfying Linear Temporal Logic (LTL) specifications. Theoretically, we establish the first PAC performance guarantees for hierarchical policies and derive bounds on abstraction quality. Experimentally, on navigation tasks with dynamic obstacles, our framework significantly improves policy generalization across rooms and enhances reliability of high-level scheduling decisions.
This study addresses the coordination challenge in decentralized inspection and maintenance planning for multi-component engineering systems by formulating the problem as a partially observable Markov decision process and employing multi-agent deep reinforcement learning (MADRL) approaches. The work encompasses a spectrum of training paradigms—from fully centralized to fully decentralized—including value decomposition and actor-critic architectures. A novel benchmark environment with tunable redundancy is introduced to systematically evaluate, for the first time, the coordination capabilities and policy optimality of various MADRL algorithms. Experimental results demonstrate that decentralized policies achieve near-optimal performance in low-redundancy series systems; however, coordination complexity increases significantly with higher redundancy. Despite this, all MADRL strategies consistently outperform an optimized heuristic baseline across tested scenarios.
To address the poor interpretability of deep reinforcement learning (DRL) policies, this paper proposes a model-agnostic interpretability distillation framework. First, it adaptively partitions the state space using Voronoi diagrams; then, it distills locally linear policies on each partition. Unlike prior approaches, it imposes no structural assumptions, thus balancing interpretability and representational capacity while preserving policy transparency and achieving performance alignment—or even improvement. The key innovation lies in coupling geometric partitioning with knowledge distillation, endowing local linear models with both theoretical traceability and empirical effectiveness. Experiments on Gridworld and classical control benchmarks demonstrate that the distilled policies exhibit clear decision logic—e.g., piecewise-linear control laws—and achieve average performance gains of 1.2%–3.7% over the original DRL policies, significantly outperforming existing interpretable baselines.
This work addresses the challenge of task offloading under stringent resource constraints in edge devices by proposing a decentralized deep reinforcement learning (DRL) agent that dynamically selects execution locations—local device, multi-access edge computing (MEC), or cloud—based on real-time conditions. The approach is the first to be deployed and evaluated on a real-world multi-device edge testbed integrated with live 5G communication, enabling a systematic assessment of the trade-offs between latency and energy consumption under local versus remote training paradigms. Experimental results demonstrate the feasibility of running DRL agents directly on end-user devices and quantitatively characterize the impact of different training deployment strategies on system performance, offering practical design guidelines for intelligent edge computing systems.
This work addresses the poor scalability and single-point-of-failure limitations of traditional centralized scheduling approaches in large-scale heterogeneous distributed systems, where dynamic workloads, resource heterogeneity, and competition for quality-of-service guarantees pose significant challenges. To overcome these issues, the authors propose DRL-MADRL, a fully decentralized multi-agent deep reinforcement learning framework that formulates task scheduling as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP). Leveraging a lightweight Actor-Critic architecture and relying solely on foundational libraries such as NumPy, the framework is efficiently deployable on resource-constrained edge devices. Experimental results on a 100-node system demonstrate that DRL-MADRL reduces average task completion time by 15.6% (30.8s vs. 36.5s), lowers energy consumption by 15.2% (745.2 vs. 878.3 kWh), and achieves an 82.3% SLA compliance rate (p < 0.001). The implementation is fully open-sourced to ensure reproducibility.
Existing multi-agent systems struggle to achieve flexible control through simple human instructions directed at only a subset of agents, and uncontrolled agents often lack the ability to autonomously infer and complement collaborative tasks. This work proposes a deep reinforcement learning–based multi-agent control framework that integrates policy following with an implicit task complementarity mechanism: by issuing commands solely to key agents, the remaining agents can adaptively infer their roles and autonomously complete the collaborative task. The approach overcomes the conventional requirement for global instructions and enables dynamic reconfiguration of collaboration structures to align with human intent. Experimental results demonstrate that the proposed method outperforms existing baselines in both collaborative flexibility and task performance.