🤖 AI Summary
This study addresses the challenges of high memory storage costs and the difficulty of distinguishing contextually recoverable from persistent information in long-horizon robotic manipulation. We propose TTT-RM, a framework that reformulates test-time training as a residual memory mechanism. Specifically, it optimizes slow weights using reconstruction residuals as the learning objective to explicitly disentangle these two information types, while leveraging fast weights to augment a bounded memory buffer for efficiently retaining irrecoverable historical states. Experimental results demonstrate that our method significantly outperforms baseline models in both simulated and real-world tasks, successfully supporting complex manipulation sequences lasting up to three minutes across eight stages.
📝 Abstract
Memory is essential for long-horizon robotic manipulation, where successful actions may depend on past events that are no longer recoverable from the current observation. As episodes grow longer, however, retaining the full history becomes increasingly costly, creating a fundamental scalability challenge for memory-augmented policies. Existing approaches address this challenge by either storing selected past observations in a bounded memory bank or compressing interaction history into a fixed-size parametric state through Test-Time Training (TTT). Yet these formulations do not explicitly distinguish between historical information that can already be recovered from the policy's current context and information that must persist beyond it. We introduce TTT-RM, which repurposes TTT as Residual Memory, using fast weights not to generically compress history but to complement a bounded memory bank by preserving task-relevant historical information that cannot be recovered from the policy's current context. Concretely, TTT-RM learns a history decoder that reconstructs historical representations from the current observation and retrieved memory. The resulting reconstruction residual captures what this context fails to explain and serves as the learning target for TTT. The TTT slow weights are optimized against this residual target so that online fast-weight updates learn to encode complementary historical information over time. The fast-weight state is then queried to produce a residual memory representation that conditions action generation. Extensive experiments on memory-intensive simulation benchmarks and real-world tasks show that TTT-RM consistently improves across multiple memory designs, outperforms diverse baselines, and supports sustained execution on a three-minute, eight-stage stowing task.