ALER: Adaptive Learnable Experience Rewriting for Reinforcement Learning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of outdated information and dynamic overwriting in memory architectures for partially observable reinforcement learning by proposing the ALER agent. The method integrates an LSTM with slot-based memory, employing a Gumbel-Softmax mechanism for adaptive discrete memory rewriting and a learned gating network to dynamically fuse multi-source experiences. Furthermore, this work formally defines memory update requirements and introduces the Rune-Mazes benchmark environment. Experimental results demonstrate that ALER achieves success rates ranging from 0.82 to 0.99 across diverse maze configurations, significantly outperforming existing baseline models and effectively enhancing decision-making robustness in complex environments.
📝 Abstract
In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the ability to keep information unchanged until it is needed. We formalize two further requirements. Rewriting sets the decision-relevant content to a value independent of the old one, and experience fusion transforms the old content by a rule that a later observation specifies. For tasks built from such updates, we count the memory states that a solution needs, and several baselines reach their lowest success rates on compositions that need more states. We introduce ALER (Adaptive Learnable Experience Rewriting), an agent that pairs an LSTM with a slot memory. An independently addressed Gumbel-Softmax write that concentrates its weight on one slot overwrites that slot, and a learned gate fuses the retrieved content with the recurrent state before the policy and value heads. We also introduce Rune-Mazes, three environments in which rune observations invert, cancel, reset, or repeat updates of a hidden cue under vector and pixel observations. Against seven baselines, ALER reaches a success rate of at least $0.82$ in all sixteen Endless T-Maze configurations and at least $0.99$ on all five Rune T-Maze compositions, and it has the highest mean success rate on four-branch Rune Multi-Corridor with an Invert rune. On pixel-based Rune MiniGrid Memory, it has a higher mean success rate than PPO-LSTM in eight of ten configurations. Project page: https://quartz-admirer.github.io/ALER-Adaptive-Learnable-Experience-Rewriting/.
Problem

Research questions and friction points this paper is trying to address.

Partially observable reinforcement learning
Memory rewriting
Experience fusion
Memory architecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Learnable Experience Rewriting
Slot Memory
Gumbel-Softmax
Partially Observable Reinforcement Learning
Experience Fusion