Reinforcement Learning with Complex (valued) Memories

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of modeling long-term dependencies and enabling effective decision-making in deep reinforcement learning under partially observable environments. It proposes replacing standard recurrent modules within the PPO architecture with unitary recurrent neural networks (uRNNs), leveraging complex vector space representations of recurrent states to enhance long-range information propagation. Specifically, three uRNN variants are designed to introduce phase degrees of freedom for optimizing gradient flow, while a phase-aware policy inspired by quantum measurement mechanisms is constructed to effectively preserve critical information from complex-valued hidden states. Evaluated on benchmark tasks including Rocksample and Craftax, the proposed method significantly outperforms existing baselines, achieving a two- to three-fold increase in cumulative reward. These results validate the effectiveness of complex-domain memory representations for intricate decision-making tasks.
πŸ“ Abstract
Partially observable environments pose a fundamental challenge in deep reinforcement learning, requiring agents to compress temporal information from observations and maintain a memory to make effective decisions. While there exist many approaches ranging from gated recurrence to attention mechanisms and model-based RL, the search for effective representational techniques that can capture long-term dependencies remains an active area of research. In this work we revisit Unitary recurrent networks (uRNNs) [Arjovsky et al., 2016, Jing et al., 2017], that demonstrated superior gradient flow and associative recall, expressing the recurrence and the hidden state in a complex vector space. Their norm preserving unitary dynamics enable information propagation through long sequences. To this end, we propose three different versions of uRNNs as drop-in replacements for recurrent PPO architectures, and demonstrate that the simple recurrence and the added degree of freedom from the phase of the complex representations enable significant gains over baselines on several memory-improvable tasks, including continuous control. We further explore how to preserve the phase information of the complex hidden state for a phase-aware policy by drawing a parallel to how quantum states are measured. With our methods reaching up to 2-3 $\times$ the reward in environments like rocksample and Craftax compared to the baselines, this work points towards an exciting new direction of representations for RL and the problem of partial observability. Code is available at: https://github.com/Sathya98/qurl
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Partially Observable Environments
Long-term Dependencies
Memory
Complex-valued Representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unitary Recurrent Networks
Complex-valued Representations
Reinforcement Learning
Partially Observable Environments
Phase-aware Policy
πŸ”Ž Similar Papers
No similar papers found.