Sparse-View 4D Gaussian Splatting via Spatiotemporal Priors and Generative Assistance

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of dynamic scene reconstruction from extremely sparse wide-baseline views captured by only six cameras. It proposes a 4D Gaussian Splatting framework that integrates region-adaptive spatial priors, motion-consistent temporal priors, and generative assistance. Key innovations include a novel foreground-background separation mechanism coupled with conditional diffusion models for virtual view completion, as well as monocular depth alignment and optical flow constraints to optimize Gaussian initialization and pseudo-supervised training. Experimental results demonstrate that the proposed method achieves a peak signal-to-noise ratio of 30.04 dB on the test set, significantly enhancing robustness under sparse data conditions. This work was awarded the overall championship in the Sparse Track at SIGGRAPH Asia 2026.
📝 Abstract
We present a 4D Gaussian Splatting framework for the Sparse-View Track of the SIGGRAPH Asia 2026 Volumetric Video Challenge, which requires dynamic scene reconstruction from only six cameras with wide baselines. To achieve robust dynamic reconstruction under such sparse views, our framework integrates three components. (1) Region-adaptive spatial priors: We use foreground masks to guide Gaussian initialization and mask voting to control densification separately for the dynamic foreground and static background. Background geometry is regularized using monocular depth aligned to metric scale. (2) Motion-consistent temporal priors: We provide supervision at intermediate times through frame interpolation and constrain projected Gaussian motion with estimated optical flow. (3) Generative assistance: We place virtual cameras in the widest angular gaps and restore their rendered images using a diffusion-based model conditioned on camera poses. The restored images are iteratively incorporated into training as pseudo-supervision. On the validation set, our framework improves full-frame PSNR from 25.60 dB for the baseline to 29.75 dB. On the official test benchmark, it achieves 30.04 dB full-frame PSNR and 27.88 dB foreground PSNR, ranking first overall in the Sparse-View Track.
Problem

Research questions and friction points this paper is trying to address.

Sparse-View
4D Gaussian Splatting
Dynamic Scene Reconstruction
Volumetric Video
Innovation

Methods, ideas, or system contributions that make the work stand out.

4D Gaussian Splatting
Sparse-View Reconstruction
Spatiotemporal Priors
Diffusion-based Generative Assistance
Dynamic Scene
🔎 Similar Papers
No similar papers found.