🤖 AI Summary
This study addresses the spatiotemporal reasoning bottlenecks of vision-language models (VLMs) in long-horizon manipulation caused by limited context windows. We propose an object-centric 4D memory architecture that maintains persistent trajectories within a shared world coordinate system, explicitly representing temporal, motion, and containment relationships. This structured memory enables VLMs to query for target selection while supporting history-agnostic low-level execution. The core innovation lies in overcoming viewpoint-dependent reasoning limitations through explicit memory mechanisms. Experimental results demonstrate a memory success rate of 96.6% and an end-to-end success rate of 88.9% in simulation, with 100% accuracy preserved under viewpoint shifts. Furthermore, real-world evaluations on a Franka robot reveal significant improvements over baseline methods across manipulation tasks.
📝 Abstract
Long-horizon manipulation often requires reasoning about state absent from the current view, such as a vanished object's location, temporal identity, or the contents of a shuffled container. We present OCC4M ("Occam"), an object-centric 4D memory that maintains persistent tracks in a shared world frame and explicitly represents temporal, motion, and containment relations. A vision-language model (VLM) queries this structured memory to select actionable targets for history-free low-level execution. Across seven simulation conditions and 350 episodes, OCC4M achieves 96.6% memory success and 88.9% end-to-end success, versus 54.6% and 57.7% for FrameSamp, a raw-history VLM baseline using Gemini 3.7 Flash with the complete observation history and the same executor. In a controlled viewpoint-transfer test, OCC4M maintains 100% memory and 98% end-to-end success after a viewpoint change, while full-history FrameSamp falls to near-zero success. On 20 fixed-camera Franka episodes, OCC4M reaches 85% joint memory accuracy, versus at most 30% for FrameSamp across context sizes from $K=16$ to the complete history, and completes 45% of full two-stage tasks. These results support explicit object-centric memory for persistent spatiotemporal reasoning in long-horizon manipulation. Qualitative videos are available at https://occ4m-sup.github.io/occ4m-supplementary/.