GrapeSplat: Geometry-Grounded Reconstruction via Amalgamated Pose-Free Encoding for Feed-Forward 3D Gaussian Splatting

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
GrapeSplat通过整合多视角线索至体素对齐场景表示,直接从学习到的网格解码高斯函数,解决了未定位图像的3D重建问题。
📝 Abstract
Feed-forward 3D Gaussian Splatting now reconstructs renderable scenes from unposed, uncalibrated images. Yet, most models supervise only photometric consistency and predict Gaussians pixel by pixel, which leaves global structure fragile and ties primitive count to image resolution and view count. To this end, GrapeSplat amalgamates multi-view cues into a voxel-aligned scene representation and decodes Gaussians directly from the learned grid, requiring no per-scene optimization or post-processing. An Atlas Encoder lifts all views into pixel-wise geometry-and-appearance features anchored at predicted 3D points. PEACH-Vox compands the unbounded scene into a bounded sparse grid through a smooth per-axis map with an exact closed-form inverse. The Sparse Decoder then consolidates the grid with sparse convolutions and decodes the full scene as multiple Gaussians per occupied cell. This amalgamated representation exploits sparse voxel occupancy, where the Gaussian count follows the occupied cells and saturates as views cover the scene, while grid resolution sets its ceiling. GrapeSplat turns unposed images into a renderable Gaussian scene in a single forward pass. Trained with 2D and 3D supervision on 8-view sequences, it generalizes zero-shot from 4 to 64 views across indoor and unbounded scenes. Code and trained weights are available at https://github.com/VAISR/GrapeSplat
Problem

Research questions and friction points this paper is trying to address.

Gaussian Splatting
unposed images
scene reconstruction
photometric consistency
sparse voxel
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometry-Grounded Reconstruction
Sparse Voxel Occupancy
Amalgamated Pose-Free Encoding
Feed-Forward 3D Gaussian Splatting
PEACH-Vox
🔎 Similar Papers
No similar papers found.
S
Si-Yu Lu
National Taiwan University
Y
Yung-Yao Chen
National Taiwan University of Science and Technology
Y
Yi Jan Chen
National Taiwan University of Science and Technology
S
Shang-Lin Li
National Taiwan University of Science and Technology
C
Ching-Chan Liao
National Taiwan University of Science and Technology
Wen-Huang Cheng
Wen-Huang Cheng
Professor, IEEE Fellow, National Taiwan University
Artificial IntelligenceMultimediaComputer VisionMachine Learning