🤖 AI Summary
This study addresses the challenge that robotic manipulation policies struggle to effectively retain and leverage historical information from demonstrations in the absence of external memory. To this end, we propose a test-time training-based memory framework that encodes observation histories into compact fast weights, which are updated online via self-supervision using solely action labels to enable efficient memory formation and retrieval. By decoupling memory and policy learning through alternating optimization, the method requires neither external reasoning models nor specialized annotations, allowing seamless integration into pretrained vision-language-action models. Evaluated on 16 RoboMME tasks, our approach increases the average success rate from 17.93% to 56.83%, significantly outperforming existing recurrent memory methods while achieving at least a threefold improvement in inference speed.
📝 Abstract
Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use historical information from action demonstrations alone, without external memory support. We introduce T$^2$Mem, a framework that develops this capability within a pretrained vision-language-action policy, without external reasoning models or memory-specific annotations. T$^2$Mem uses test-time training to encode observation history into compact fast weights through online self-supervised updates, avoiding repeated processing of the full history. An observation-grounded interface extracts vision-language information for memory formation and supplies retrieved context to the action expert. Action supervision shapes what the memory learns to retain and use, while alternating memory-policy learning gives each component a fixed counterpart during optimization. Across 16 RoboMME tasks, T$^2$Mem improves average success from 17.93% to 56.83% over the memory-free base policy and outperforms the recurrent-memory methods reported in the benchmark, while controlled profiling indicates at least 3x inference speedup over explicit methods. Project website: https://yzliu84.github.io/T2MEM-project/