π€ AI Summary
This study addresses the challenge of dynamic 3D reconstruction from monocular static-camera videos, where multi-view supervision is inherently unavailable. To overcome this limitation, we propose a depth-guided proxy image enhancement method that leverages a foundational monocular depth estimation network to synthesize multi-view supervisory signals. By integrating 3D Gaussian Splatting with a deformation network to model temporal dynamics, our approach pioneers the use of proxy image synthesis to circumvent the absence of multi-view constraints. Experiments on the DyNeRF dataset demonstrate that the proposed method achieves high-quality novel view synthesis relying solely on monocular depth priors. Notably, it significantly outperforms existing approaches that depend on strong priors such as scene flow, establishing a new state-of-the-art for dynamic scene reconstruction under monocular settings.
π Abstract
We present ProDyGS, a novel dynamic 3D Gaussian Splatting framework for high-quality novel view synthesis from videos captured by a single static camera. While existing methods rely on multi-view setups or significant camera motion for geometric constraints, our approach addresses the challenging scenario where multi-view supervision is completely absent. We overcome this limitation by generating synthetic multi-view supervision through depth-guided proxy image synthesis. Specifically, we estimate temporally consistent depth maps using foundational monocular depth networks, then construct 3D Gaussian representations that generate proxy images from arbitrary viewpoints. A deformation network learns temporal dynamics by warping canonical Gaussians using this augmented supervision. Experiments on the DyNeRF dataset demonstrate that our method achieves state-of-the-art performance while requiring only monocular depth estimation as external supervision, outperforming approaches that rely on stronger priors such as scene flow.