🤖 AI Summary
This study addresses the limitation of vision-language-action models in long-horizon manipulation, where a single internal memory fails to span the full spatiotemporal spectrum. To this end, we propose Ledger, an architecture grounded in the principle that memory type dictates storage location: short-term perceptual memory is maintained within the policy, while long-term object memory is externalized into an explicit ledger. By integrating SAM3 tracking, VLM captioning, and LLM planning, Ledger dynamically selects memory sources without task-level routing. When integrated with the π0.5 policy, our approach achieves a state-of-the-art average score of 64.3% across four suites on the RoboMME benchmark, significantly outperforming existing methods and substantially improving performance in object referencing and permanence tasks.
📝 Abstract
Memory is essential for long-horizon, partially observed robotic manipulation: a robot must remember which object was placed in a drawer, whose cup it moved, or how many action cycles have elapsed. Recent vision-language-action (VLA) models embed memory directly inside the policy, but benchmarks show no single in-policy mechanism covers all spatio-temporal dimensions, trailing oracle methods by a wide margin. We argue that memory type dictates where memory should reside: short-term perceptual memory (repetition, timing, retracing) belongs inside the policy, while long-term object memory (persistent spatial state, containment, event history) belongs outside as an explicit, readable record. We present Ledger, a harness that realizes this split over a single fine-tuned $π_{0.5}$ policy by pairing an in-policy frame-sampling memory with an external spatio-temporal object memory, the ledger, built from a SAM3 tracker and a VLM captioner of the demonstration and read by an LLM planner that decides at step boundaries. On RoboMME, Ledger reaches the highest four-suite average among the evaluated methods, 64.3% (vs. 45.9% for the strongest prior method under identical evaluation), leading object reference (60.7% vs. 40.3%) and object permanence (86.7% vs. 56.2%) using a single set of weights. Choosing the memory source at runtime, from the instruction and the record, removes the need for a task-level router.