🤖 AI Summary
To address the inefficiency and poor temporal consistency in long-term Gaussian scene reconstruction caused by frequent object mutations and subtle changes in everyday environments, this paper proposes a templated temporal Gaussian modeling framework. Methodologically, it leverages structured Gaussian object templates as transferable priors, integrated with sparse-view observations and few-shot self-adaptive optimization, enabling lightweight, incremental scene evolution modeling via spatio-temporal transformations; initialization employs Gaussian splatting for rapid online updates. Experiments on a newly constructed real-world dynamic scene dataset demonstrate significant improvements in reconstruction fidelity and inter-frame geometric consistency, alongside a 3.2× speedup in inference latency and a 57% reduction in memory overhead. To our knowledge, this is the first approach achieving high-fidelity, long-duration, low-latency Gaussian scene chronology construction.
📝 Abstract
Recent advances in novel-view synthesis can create the photo-realistic visualization of real-world environments from conventional camera captures. However, acquiring everyday environments from casual captures faces challenges due to frequent scene changes, which require dense observations both spatially and temporally. We propose long-term Gaussian scene chronology from sparse-view updates, coined LTGS, an efficient scene representation that can embrace everyday changes from highly under-constrained casual captures. Given an incomplete and unstructured Gaussian splatting representation obtained from an initial set of input images, we robustly model the long-term chronology of the scene despite abrupt movements and subtle environmental variations. We construct objects as template Gaussians, which serve as structural, reusable priors for shared object tracks. Then, the object templates undergo a further refinement pipeline that modulates the priors to adapt to temporally varying environments based on few-shot observations. Once trained, our framework is generalizable across multiple time steps through simple transformations, significantly enhancing the scalability for a temporal evolution of 3D environments. As existing datasets do not explicitly represent the long-term real-world changes with a sparse capture setup, we collect real-world datasets to evaluate the practicality of our pipeline. Experiments demonstrate that our framework achieves superior reconstruction quality compared to other baselines while enabling fast and light-weight updates.