🤖 AI Summary
This study addresses the challenge of maintaining persistent scene states in streaming 3D reconstruction to enable online fusion and real-time rendering. To this end, it proposes a feedforward framework based on adaptive-sized persistent tokens. Specifically, a spatial-aware Transformer is employed to integrate new observations with historical states, while a selective admission module distinguishes between scene updates and expansions. Furthermore, a hierarchical decoder directly generates 3D Gaussians without requiring the caching of historical frames. Evaluated across four benchmarks, the proposed method achieves competitive streaming rendering quality through a compact Gaussian representation, effectively resolving the memory and efficiency bottlenecks inherent in online 3D reconstruction.
📝 Abstract
Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive. Latent spatial tokens offer a promising representation for this purpose, but constructing them from an image collection leaves open how to maintain them online, where each observation may both revisit known regions and reveal new content. We introduce S2Tok, a feed-forward framework that maintains a size-adaptive, persistent scene state from uncalibrated image streams. Its central idea is to distinguish updates to the existing representation from selective expansion. A spatially informed transformer integrates each incoming observation with the persistent scene tokens, while a learned admission module selectively expands the representation to limit redundant storage. A hierarchical decoder and Gaussian head convert the evolving state into non-pixel-aligned 3D Gaussians, enabling novel-view rendering without caching previous frames. Experiments across four benchmarks demonstrate competitive streaming rendering quality with compact Gaussian representations. These results support latent spatial tokens as a persistent computational state for online 3D reconstruction, combining learned scene updates with explicit Gaussian rendering.