Slot3R: Set-Associative Spatial Memory for Streaming 3D Reconstruction

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of evidence loss in streaming 3D reconstruction caused by spatial memory conflating locations with states. To this end, we propose a set-associative spatial memory mechanism that utilizes locations solely for addressing, thereby permitting multiple states to coexist under shared addresses and preserving complementary observations. Furthermore, by freezing the Point3R pretrained backbone and decoupling memory storage from access via bounded sparse readout, the method enables efficient online processing of expanding scenes. Experimental results demonstrate that point cloud accuracy errors are reduced by over 57% alongside significant decreases in trajectory error, while maintaining approximately 19 FPS on long sequences without incurring memory overflow.
📝 Abstract
Streaming 3D reconstruction must preserve evidence from each frame while processing an expanding scene online. Spatial memory is a natural fit because it organizes history by reconstructed 3D location. Yet Point3R uses spatial proximity both to associate a new observation with an existing memory entry and to decide whether to fuse it, conflating co-location with state identity. Because pointers summarize image patches, nearby pointers may encode distinct surfaces, viewpoints, or visibility conditions; averaging them can destroy complementary evidence before later frames disambiguate it. We argue that location should determine address, not whether observations must merge. Slot3R realizes this principle as a training-free, set-associative retrofit that lets multiple states coexist at a shared address while keeping the pretrained Point3R backbone frozen. A bounded sparse readout further decouples persistent storage from per-frame decoder access. At 300-500 sampled frames, Slot3R reduces Point3R's point-cloud accuracy error (Acc) by 57.1%-63.1% on 7Scenes and 64.0%-72.0% on NeuralRGBD, lowers Sim(3)-aligned absolute trajectory error (ATE) on all three pose benchmarks, and remains competitive on video-depth estimation. It completes all evaluated settings from 600 to 1000 sampled frames at about 19 FPS under the same protocol, whereas Point3R and InfiniteVGGT run out of memory at 800 frames and beyond.
Problem

Research questions and friction points this paper is trying to address.

streaming 3D reconstruction
spatial memory
state identity
evidence preservation
memory efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Set-Associative Spatial Memory
Streaming 3D Reconstruction
Training-free Retrofit
Bounded Sparse Readout
State Disambiguation
🔎 Similar Papers
2024-09-14IEEE Robotics and Automation LettersCitations: 0