RDGSplat: Render-Dedicated Geometry for Novel View Synthesis

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient rendering quality of 3D foundation models and the degradation of metric predictions caused by fine-tuning. To this end, we propose RDGSplat, a framework that introduces the first rendering-specific geometry decoding mechanism to decouple metric estimation from rendering functionalities. By incorporating a target-pose-conditioned adapter, our method optimizes novel view synthesis via photometric supervision while keeping the backbone weights frozen. Experimental results on the RE10K benchmark demonstrate that RDGSplat improves PSNR to 24.266 dB while fully preserving depth and pose prediction performance. With only 205.5M additional parameters, the proposed framework achieves efficient, lossless rendering enhancement for 3D foundation models.
📝 Abstract
3D foundation models enable efficient novel view synthesis by carrying a Gaussian head on the representation they already use for reconstruction. However, the views they render fall short of the geometry they recover, because that geometry is estimated under a metric objective and never scored on how it renders. Recent methods alleviate this by updating the backbone weights, but they thereby discard the metric predictions the model was built for and must be repeated for every new backbone. To this end, we propose RDGSplat, a framework that decodes a second geometry dedicated to rendering from a frozen 3D foundation model, leaving its metric predictions intact. In particular, we devise Render-Dedicated Geometry Decoding, which duplicates the pretrained decoders and optimizes the duplicates under photometric supervision alone. Then, a Target-Pose Conditioned Adapter is introduced to reformulate the representation those decoders read, conditioned on the target camera pose rather than the target image. Extensive experiments show that RDGSplat improves novel view synthesis across three feed-forward backbones on four benchmarks, with every pretrained weight frozen. On RE10K, it raises WM2.0 from 20.918 to 24.266\,dB while training 205.5\,M added parameters against a frozen 1.4\,B backbone, and the depth and pose the same model predicts are unchanged.
Problem

Research questions and friction points this paper is trying to address.

Novel View Synthesis
3D Foundation Models
Render-Dedicated Geometry
Metric Prediction Preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Novel View Synthesis
3D Foundation Models
Render-Dedicated Geometry Decoding
Target-Pose Conditioned Adapter
Gaussian Splatting
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhijie Zheng
University of California, Davis
X
Xinhao Xiang
University of California, Davis
Jiawei Zhang
Jiawei Zhang
University of California, Davis (UC Davis)
Machine LearningFoundation Model