ByteSplat: Efficient Distributed 3D Gaussian Splatting Training via Intra- and Inter-GPU communication reduction

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory access and inter-GPU communication bottlenecks encountered in distributed training of large-scale 3D Gaussian Splatting (3DGS) by proposing a hardware-software co-optimization framework. Methodologically, it introduces the first single-kernel fusion of forward and backward rasterization to minimize on-chip data movement, designs a hardware-aware adaptive pruning strategy governed by shared memory constraints, and incorporates gradient sparse-coded aggregation to compress inter-chip communication volume. Experimental evaluations demonstrate that the proposed approach achieves up to a 6.1× training speedup in an eight-GPU configuration while significantly reducing data traffic and preserving reconstruction quality.
📝 Abstract
3D Gaussian Splatting (3DGS) enables photorealistic scene reconstruction, but training large-scale scenes requires substantial memory and computation. Distributing training across multiple GPUs increases available memory capacity, yet its performance is strictly constrained by data movement. We identify two dominant bottlenecks: repeated off-chip memory accesses during forward and backward rasterization, and inter-GPU communication of partial Gaussian gradients. To address these bottlenecks, we present ByteSplat, a distributed 3DGS training framework that jointly reduces intra- and inter-GPU data movement. First, ByteSplat fuses forward rasterization and backward rasterization into a single GPU kernel, retaining intermediate results on-chip to eliminate redundant off-chip transfers. Second, to alleviate the increased on-chip storage by fused rasterization, ByteSplat introduces hardware-aware pruning that considers both rendering quality and per-tile shared-memory constraints, increasing the overall performance of fused execution. Lastly, ByteSplat exploits gradient sparsity to eliminate zero partial gradients from inter-GPU communication. GPU-efficient encoder and decoder kernels compact the remaining records and directly aggregate the received gradients on their owner GPUs, reducing both communication volume and local aggregation overhead. We evaluate ByteSplat across six datasets. Compared with the baseline, ByteSplat reduces intra-GPU off-chip traffic and backward inter-GPU communication volume by 63.4% and 65.8%, respectively. On eight GPUs, ByteSplat achieves up to 6.1$\times$ training speedup while preserving reconstruction quality.
Problem

Research questions and friction points this paper is trying to address.

3D Gaussian Splatting
Distributed Training
Intra-GPU Communication
Inter-GPU Communication
Memory Bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D Gaussian Splatting
Distributed Training
Kernel Fusion
Gradient Sparsity
Communication Reduction
🔎 Similar Papers
2024-06-26International Conference on Learning RepresentationsCitations: 30