🤖 AI Summary
This study addresses the challenges of disparity extrapolation and occlusion rendering in feed-forward 3D Gaussian Splatting from stereo cameras, proposing an optimization-free method for rapid scene reconstruction. The approach reuses intermediate features from a frozen pre-trained stereo matching network, combined with calibrated disparity to anchor geometry for predicting metric-scale 3D Gaussians. To effectively handle occluded content, it introduces a dual-layer Gaussian mechanism alongside an extended canvas to enhance scene capacity. Furthermore, a high-quality synthetic dataset, SceneSplat-Stereo, is constructed. Experimental results demonstrate that the proposed method significantly outperforms existing strong baselines on novel view synthesis tasks across both real-world and photorealistic stereo benchmarks, validating its effectiveness.
📝 Abstract
Feed-forward 3D Gaussian Splatting (3DGS) enables reconstruction without per- scene optimisation, but practical stereo-camera applications require nearby-view extrapolation beyond the input views. Stereo depth anchors visible surfaces, yet rendering newly exposed regions also requires learned appearance and additional scene capacity. We introduce StereoGaussians, which predicts a metric 3DGS representation from a single calibrated stereo pair. It reuses intermediate repre- sentations from frozen pretrained stereo networks to predict Gaussian attributes, while calibrated disparity anchors the geometry. A second Gaussian layer and an expanded image canvas provide capacity for disoccluded and outside-field-of- view content. For training, we construct SceneSplat-Stereo from quality-filtered 3DGS teachers, pairing stereo inputs with nearby target views across 803 training scenes. Experiments on unseen real and photorealistic stereo benchmarks demon- strate improvements over strong view-synthesis baselines, while ablation studies support our main design choices.