Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses how embodied agents can construct persistent 3D object memories from egocentric videos to accurately recall object locations and contextual information after they leave the field of view. To this end, it proposes Ledger, a system that employs spatiotemporal clustering to mitigate localization noise and associate cross-temporal observations, while leveraging repeated evidence to track object movements. Furthermore, Ledger utilizes natural language generation to produce concise descriptions for preserving fine-grained details, thereby constructing a persistent 3D scene graph retrievable without accessing the raw video. Experimental results demonstrate that the system significantly enhances memory robustness, achieving accuracies of 42.6% and 38.5% on HD-EPIC and UCS-Bench, respectively, with a median error of only 0.99 meters on Ego4D. These findings validate the complementary advantages of temporal persistence and contextual retrieval.
📝 Abstract
As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later. We are capable of recalling where we left something or what was inside a container, even without knowing we would need it later. Here, we study how an embodied assistant can build a similar memory from egocentric videos, by observing a person's day-to-day activities. We present Ledger, a persistent 3D object memory that combines object locations, their histories, and contextual descriptions. It associates observations across the recording and retains objects after they leave the view, including those the person never touches. It clusters each object's observations by resting locations and records a move only after repeated evidence, reducing the effect of localization noise. Short descriptions preserve details such as an object's contents or supporting surface. It saves these records to later answer spatial questions without having to access the original images or video. Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5% and localizes Ego4D objects with a 0.99 m median error on returned predictions. Our analyses identify complementary roles for temporal persistence, contextual descriptions, and retrieval. Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction. Per-scene construction partially recovers the performance lost across scene changes compared to that of single scene streams.
Problem

Research questions and friction points this paper is trying to address.

Egocentric video
3D object memory
Embodied assistant
Spatial reasoning
Persistent memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Persistent 3D Object Memory
Egocentric Videos
Spatial Reasoning
Noise Reduction
Contextual Descriptions
🔎 Similar Papers
No similar papers found.