🤖 AI Summary
This study addresses the lack of long-term memory in robotic manipulation policies and the limitations of existing methods that increase inference costs or restrict generalizability. To this end, we propose a plug-and-play long-term memory module. This module implements fixed-capacity associative memory via gated Delta rule linear attention and a parallel token update mechanism, utilizing an architecture-agnostic readout layer to adapt to arbitrary pretrained policies. Integration requires neither backbone modification nor model retraining, relying solely on lightweight post-training. Experiments demonstrate that, with the pretrained policy frozen, our approach significantly outperforms short-context baselines and existing compact memory schemes, efficiently endowing models with robust long-term memory capabilities.
📝 Abstract
Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a short temporal window. Equipping such policies with long-term memory remains challenging: existing approaches either feed the backbone multi-frame observation windows, which substantially increase inference cost, or rely on pre-defined semantic features, which limit task generality and may also require the retraining of the backbone to adapt to the memory. We introduce DRAM (Delta-rule Recurrent Associative Memory), a plug-and-play memory module that can be attached to a wide range of pretrained robotic policies, endowing them with long-horizon memory without architectural modification or backbone retraining, requiring only task-specific post-training of the memory module and action expert. DRAM maintains a fixed-size associative memory using gated delta-rule linear attention, with a modified update that incorporates all tokens within each frame in parallel. An architecture-agnostic readout integrates historical context into action prediction across different policy architectures. Experiments show that DRAM consistently improves frozen pretrained policies over short-context baselines and alternative compact memory designs, validating its effectiveness as a fixed-size, post-hoc memory module trained with the backbone frozen.