DensiTok: Making Feed-Forward 3D Gaussian Splatting See More Views Than It Is Given

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation in reconstruction quality of feed-forward 3D Gaussian Splatting under sparse-view settings caused by representational deficiencies. To this end, we propose DensiTok, a module that directly densifies geometric tokens within a low-dimensional latent space. By compressing, completing, and decoding these tokens, DensiTok effectively injects information from unobserved views without requiring image synthesis or additional encoders, while keeping the backbone network frozen. Coupled with a flow matching training strategy, our method significantly enhances few-view reconstruction performance across multiple benchmarks, substantially narrowing the accuracy gap between sparse-view and dense-view reconstructions.
📝 Abstract
Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene in a single forward pass, replacing per-scene optimization with a network trained across many scenes. Its quality, however, degrades sharply as the number of input images drops. The bottleneck is upstream of the reconstruction heads: from a few unposed views, the internal representation they read carries no evidence for unobserved regions, leaving holes, floaters, and blur. The common remedy supplies that evidence as pixels, synthesizing extra views with an image or video generator and re-encoding them, which is costly and not 3D-consistent by construction. We instead densify the evidence itself. We present DensiTok, a plug-in module for pretrained feed-forward 3DGS models that densifies their internal geometry tokens directly, making a frozen backbone behave as though it had observed many more views than it was given. DensiTok compresses those tokens into a compact latent space, completes the latents of the unobserved viewpoints in a single flow-matching step conditioned on camera geometry, and decodes them back into tokens that the original reconstruction heads. The same module design can be integrated into different pretrained predictors while keeping each backbone and its reconstruction heads frozen. Completion in a low-dimensional latent space requires no image synthesis or additional encoder passes. Across three pretrained backbones and two benchmarks, DensiTok consistently improves sparse-view reconstruction and recovers much of the gap to dense-view reconstruction.
Problem

Research questions and friction points this paper is trying to address.

Feed-forward 3D Gaussian Splatting
Sparse-view reconstruction
Geometry token densification
Novel view completion
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D Gaussian Splatting
Feed-Forward Reconstruction
Flow Matching
Sparse-View
Plug-and-Play
🔎 Similar Papers
No similar papers found.