LoCoSplat: Real-Time Feed-Forward 3D Gaussian Splatting with Minimal 3D Reasoning

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the slow inference and high memory consumption of feed-forward 3D Gaussian Splatting methods that rely on heavy 3D networks. We propose LoCoSplat, which eliminates complex 3D reasoning by substituting global modeling with local context. Specifically, it derives Gaussian attributes through dual-resolution grid splatting, 16-dimensional linear projections, and pointwise micro-MLPs containing only 0.14M parameters for feature aggregation. The entire encoding pipeline is executed within a single CUDA graph, yielding a minimalist yet highly efficient architecture. On the RealEstate10K dataset, LoCoSplat achieves state-of-the-art performance across all metrics. For 6-view reconstruction, it requires merely 33 milliseconds, delivering a 4.2× inference speedup, a 2.7× training acceleration, and a 6.7× reduction in GPU memory usage compared to existing state-of-the-art approaches.
📝 Abstract
Feed-forward 3D Gaussian Splatting (3DGS) increasingly aggregates multi-view evidence with heavy learned 3D networks. We propose LoCoSplat (Local-Context Splatting), motivated by the observation that a Gaussian is a local primitive: once depth is predicted, what the 3D stage must add (scale, rotation, opacity) depends on the point cloud around each anchor, and a fixed local average of that neighbourhood is enough to supply it, no heavy network required. LoCoSplat realises exactly this average: it splats a 16-d linear projection of the point features into a fine and a coarse grid and reads both back at each anchor with a 0.14M-parameter pointwise MLP; with no learned 3D network and no dynamic sparse computation, its whole encoder runs as one fp16 CUDA graph. On RealEstate10K, LoCoSplat outperforms every prior feed-forward method on PSNR, SSIM, and LPIPS at 6, 12, and 24 views, with a margin that widens as views densify (+3.3 PSNR over VolSplat, the prior voxel-aligned state of the art, at 24 views) and grows further under zero-shot transfer to ACID and fine-tuning on ScanNet. It reconstructs a 6-view scene in 33 ms on one NVIDIA RTX PRO 6000 GPU, the fastest of seven feed-forward methods and $4.2\times$ faster than the previous state of the art, trains $2.7\times$ faster ($5.9\times$ at 24 views), and uses $6.7\times$ less inference memory.
Problem

Research questions and friction points this paper is trying to address.

3D Gaussian Splatting
feed-forward reconstruction
real-time rendering
computational efficiency
multi-view aggregation
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D Gaussian Splatting
Feed-Forward
Local Context
CUDA Graph
Pointwise MLP
🔎 Similar Papers
No similar papers found.