Score
Design and implement observation-augmentation methods that attach age-of-information (AOI) or delay metadata to agent observations, producing delay-aware state representations for decentralized partially observable decision processes. Build and analyze coordination and recurrent-policy inputs that consume these augmented observations to enable delay-aware decision making, preserve triangulation validity under latency, and measure impacts on policy performance.
Time delays compromise the Markov property of control systems, leading to degraded reinforcement learning performance and heightened stability risks. This work presents the first systematic survey and classification of five categories of reinforcement learning approaches designed to address perception, actuation, and communication delays: state augmentation, recurrent policies, predictor-augmented modeling, robust domain randomization, and safety-constrained reinforcement learning. By framing these methods within a unified perspective, the study compares their applicability and inherent trade-offs, offering practical guidelines for method selection. Furthermore, it identifies critical open challenges—such as stability certification and multi-agent coordination—and outlines promising directions for future research, thereby providing a theoretical foundation for designing reliable controllers in delay-prone environments.
In real-world multi-agent systems, asynchronous and stochastic observation delays are prevalent, causing distorted local observations and severely degrading policy learning performance. To address this, we formally introduce the Decentralized Stochastic Individual Delay Partially Observable Markov Decision Process (DSID-POMDP) — the first model to rigorously characterize agent-specific, random observation delays in decentralized settings. Building upon it, we propose Rainbow Delay Compensation (RDC), an end-to-end training framework integrating a delay-aware encoder, a temporal alignment module, and a rainbow-style variant of Q-learning, enabling robust policy learning and cross-delay generalization under heterogeneous delay patterns. Evaluated on MPE and SMAC benchmarks, RDC significantly mitigates delay-induced performance degradation, restoring near-no-delay performance across diverse stochastic delay distributions. Our results demonstrate both effectiveness and strong generalization capability across unseen delay conditions.
We formally establish the equivalence between Observation Delay (OD) and Action Delay (AD) in cooperative partially observable multi-agent systems using observation-action histories. We show that both systems generate identical admissible joint-policy sets, and their induced state-action-observation trajectories are identical in distribution, leading to identical optimal solutions in Decentralized Partially Observable Markov Decision Processes (Dec-POMDPs). This formally generalizes existing infinite-horizon single-agent results to any-horizon partially observable cooperative multi-agent problems with decentralized policy execution, and allows any mixed-delay configuration to be reduced to a pure OD system. We further prove that in Transition-Independent MDPs (TI-MDPs), the observation-action history reduces to a tractable minimal local augmented state. However, we show through numerical experiments that although the optimal solution spaces are structurally isomorphic, the practical learning dynamics are fundamentally different. First, using the minimal local augmented state, the equivalence no longer holds when transitions are not independent. Second, operational constraints and causal credit-assignment errors in Temporal Difference (TD) algorithms induce different learning behaviors across regimes. Finally, leveraging this structural equivalence to bypass these learning challenges, we demonstrate successful multi-agent zero-shot policy transfer from OD to AD, paving the way for unified, efficient solution methods in complex delayed systems.
This work addresses the significant performance degradation commonly observed in real-world multi-agent reinforcement learning (MARL) systems due to observation and communication delays as well as packet loss. To mitigate these issues, the authors propose a modular, plug-and-play state estimation layer that replaces delayed observations with belief states during execution. This layer integrates a learned gated dynamic model with a recursive Kalman filter to enable robust estimation of instantaneous states under asynchronous and incomplete measurements. Notably, the approach requires no modifications to the underlying MARL algorithm, architecture, or reward function. Experimental results demonstrate that the method substantially enhances policy robustness against communication delays and packet loss across multiple multi-agent continuous control benchmarks, with particularly strong performance in scenarios requiring tight coordination or exhibiting dynamic instability.
Real-world sensor observations often suffer from stochastic delays and out-of-order arrivals, whereas standard reinforcement learning assumes instantaneous observations; existing approaches inadequately model such delays within the partially observable Markov decision process (POMDP) framework. This paper formally models stochastic observation delay within POMDPs for the first time and introduces a sequential belief-state update mechanism that dynamically fuses delayed, out-of-order observations—overcoming the limitations of conventional history-stacking methods. The mechanism is robust to shifts in delay distribution and integrates seamlessly into model-based RL frameworks such as Dreamer. Experiments on simulated robotic control tasks demonstrate that our method significantly outperforms existing baselines and heuristic strategies across diverse delay patterns—including stochastic, bursty, and heavy-tailed delays—exhibiting strong effectiveness and generalization capability.
This work addresses the challenge of degraded information timeliness and action misalignment in multi-agent collaboration caused by cross-timestep communication delays. It formalizes delayed communication as a Decentralized Communication Partially Observable Markov Game (DeComm-POMG) and introduces a Communication Gain and Delay Cost (CGDC) decomposition mechanism. Building upon this, the authors propose an adaptive communication strategy that triggers message exchange only when the net benefit is positive. They develop CDCMA, an actor-critic framework integrating future observation prediction, CGDC-guided attention, and dynamic message requesting to enhance information fusion efficiency. Experimental results demonstrate that the proposed method significantly outperforms baseline approaches across diverse delay settings in Cooperative Navigation, Predator-Prey, and SMAC tasks, achieving superior performance, robustness, and generalization capability.
This work addresses the degradation in 3D localization accuracy in anti-drone systems caused by existing methods’ neglect of cumulative delays in detection, communication, and decision propagation within multi-agent active visual triangulation. To mitigate this, the authors propose a delay-aware, uncertainty-driven multi-agent reinforcement learning framework that enhances observation modeling in decentralized partially observable Markov decision processes (Dec-POMDPs) using Age of Information (AoI), introduces a perception-consistent reward mechanism, and—uniquely—integrates multi-source uncertainties from pixels, poses, gimbals, and intrinsic camera parameters into covariance propagation. Experiments in 4,096 parallel environments demonstrate a root-mean-square error of 0.547 ± 0.217 meters and 78.1% triangulation validity, outperforming angle-noise-only models by a 2.8× reduction in error and vastly exceeding MLP-based policies (validity <0.7%), thereby confirming the critical role of recurrent memory in compensating for system delays.
This work addresses the challenges of unreliable status updates and heterogeneous operational costs in integrated sensing and communication systems by jointly optimizing the Age of Information (AoI) and long-term aggregate cost. In the single-source setting, a Markov decision process is formulated, revealing an optimal policy with a monotone threshold structure, and a state-space truncation method with rigorous error bounds is developed. For the multi-source scenario, the problem is cast as a restless multi-armed bandit, leading to a broadly applicable approximate Whittle index scheduling policy. Theoretical analysis provides guarantees on both policy structure and truncation error, while numerical experiments demonstrate that the proposed approximate Whittle index significantly outperforms baseline methods in both indexable and non-indexable regimes.
This work addresses the significant performance degradation of reinforcement learning policies under observation delays, a common challenge in real-world settings. In stochastic Markov decision processes, delayed observations inherently diverge from the true current state, introducing bias that impairs policy effectiveness. The paper presents the first theoretical characterization of this bias and introduces a novel delay-aware policy optimization framework. This framework explicitly models the mapping from delayed observations to the current true state using a diffusion model and incorporates an uncertainty-aware mechanism to reweight and correct policy updates. Evaluated on a range of continuous robotic control tasks with stochastic and long observation delays, the proposed method substantially outperforms existing approaches, demonstrating both strong robustness and superior performance.
This work addresses the challenge of enabling agents to implicitly convey internal state information through their actions in communication-constrained environments, thereby facilitating accurate external observation. The authors propose a method that directly embeds state observability into the reinforcement learning reward function, guiding the policy to actively expose informative state signals while preserving primary task performance. By integrating reinforcement learning with observability-aware optimization, the approach successfully trains control policies with high observability in an aerial tracking task. Experimental results demonstrate that the resulting policies significantly enhance the accuracy of state reconstruction by external observers, with negligible degradation to the main task performance.