🤖 AI Summary
This study addresses the over-reliance of robotic manipulation policies on historical visual, counting, and temporal information by proposing ReCAT, a language-conditioned policy. The method leverages structured recurrent memory to integrate multimodal observation streams, combining Mamba-2 with causal attention layers to enhance decision-making. It introduces a novel block-wise cross-attention memory readout mechanism, revealing the distinct advantages of different update rules in spatial recall and timing tasks. The architecture integrates instruction encoding, Mamba-2 recurrent layers, and a flow-matching Transformer decoder. Experiments demonstrate that ReCAT achieves success rates of 95.3% and 62.4% on the LIBERO and RMBench benchmarks, respectively. Furthermore, it attains an average success rate of 66.7% on real-world robots, substantially outperforming the baseline at 8.3%.
📝 Abstract
Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors, such as recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time. We present ReCAT, a language-conditioned policy with structured recurrent memory. An instruction-conditioned encoder forms features from the current observation. A recurrent memory integrates the observation stream through Mamba-2 layers and one causal attention layer. A flow-matching Transformer decoder reads the current and the historical representation through separate cross-attention in every block. ReCAT reaches 95.3\% average success on LIBERO and 62.4\% on RMBench, with the best or tied-best result on six of nine tasks. On three real-robot tasks probing spatial recall, event counting, and interval timing, the best ReCAT variant reaches 66.7\% average success, against 8.3\% for the strongest short-history baseline. Controlled comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance. They also show that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule updates on spatial recall. Project website is at https://intuitive-robots.github.io/ReCAT