Simple Agentic Memory for Generalist Robot Policies

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that missing interaction-derived states, such as identity and progress, in robot control frequently lead to failures in memory-dependent tasks. To overcome this limitation, this work proposes a state-based rather than visual-history-based robotic memory paradigm, constructing a training-free memory-layer agent. The architecture leverages frozen perception tools to maintain compact, typed states, integrating structured retrieval with multimodal grounding mechanisms to enable efficient state access and alignment with current views. Evaluated on 16 tasks within the RoboMME benchmark, the proposed approach achieves an average success rate of 67.17%, significantly outperforming the strongest baseline at 44.51%. Furthermore, ablation studies targeting relational and referential states validate the effectiveness of each individual module.
📝 Abstract
Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual frame may explicitly represent, such as persistent identity relations, accumulated progress, or ordered procedures. We introduce Simple Agentic Robot Memory (SimpleARM), a training-free memory layer for frozen generalist robot policies. From the task instruction, SimpleARM specifies what to monitor; frozen perceptual tools maintain compact typed state online; structured access retrieves that state only when a proposed subgoal depends on history; and current-view grounding resolves recalled entities before execution. We evaluate SimpleARM on RoboMME, a benchmark of memory-dependent robot manipulation tasks that require history information no longer available in the current observation. Across all 16 tasks and three policy seeds, SimpleARM achieves 67.17% mean success, compared with 44.51% for the strongest non-oracle baseline. Matched ablations show mechanism specificity: removing relation, reference, progress, or route state produces large losses where the affected state is retrieved for control, while largely sparing other tasks. These results support a state-based view of robot memory: effective memory for control is not simply retained visual history, but compact task-relevant state derived from the interaction history.
Problem

Research questions and friction points this paper is trying to address.

robot memory
generalist robot policies
visual-memory systems
interaction-derived state
robot manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Memory
Generalist Robot Policies
Training-free
Structured State
RoboMME