DeltaSplat: Iterative Gaussian Refinement for Pose-Free Feed-Forward 3D Gaussian Splatting

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the geometric and photometric misalignment caused by camera estimation errors in pose-free feed-forward 3D Gaussian Splatting. To mitigate this issue, we propose a lightweight iterative refinement module that leverages Plücker rays and depth priors to resolve the underdetermined problem. Specifically, the method predicts Gaussian updates for error correction through a dual-branch convolutional mixer, per-pixel ray encoding, and residual rendering feedback. This module introduces only approximately 2.2% additional parameters while preserving fully feed-forward inference. Evaluated on the DL3DV dataset, our approach achieves a PSNR of 26.64 dB, outperforming the baseline by 1.75 dB and even surpassing benchmark performance obtained using ground-truth cameras.
📝 Abstract
Pose-free feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene from sparse, unposed images in a single network pass, removing the need for camera calibration and per-scene optimization. However, camera estimation errors propagate into the predicted Gaussians and compound the geometric and photometric inaccuracies of single-pass prediction. To correct these errors, we introduce DeltaSplat, a lightweight Gaussian refinement module for pose-free feed-forward 3DGS. It iteratively renders the current Gaussians at the input context views and predicts per-Gaussian updates from the resulting residuals. A 2D residual alone, however, underdetermines the 3D correction. DeltaSplat therefore conditions each update on per-pixel Plücker rays and rendered depth as a soft geometric prior. A dual-branch convolutional mixer efficiently encodes these inputs, and per-attribute heads decode the fused features into position, opacity, and color updates. The module adds only ~2.2% parameters to the backbone and remains fully feed-forward at inference. On DL3DV, DeltaSplat reaches 26.64 dB PSNR in the pose-free setting, improving its state-of-the-art backbone by 1.75 dB and surpassing even baselines supplied with ground-truth cameras; consistent gains hold across 6-24 views and all camera regimes.
Problem

Research questions and friction points this paper is trying to address.

pose-free 3D Gaussian Splatting
camera estimation errors
feed-forward reconstruction
geometric inaccuracies
photometric inaccuracies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pose-Free 3D Gaussian Splatting
Iterative Gaussian Refinement
Plücker Rays
Dual-Branch Convolutional Mixer
Feed-Forward Reconstruction
🔎 Similar Papers
No similar papers found.