🤖 AI Summary
This study addresses the challenge that 3D Gaussian Splatting (3DGS) model sizes fail to adapt to scene scales, resulting in detail loss for large scenes or computational redundancy for small ones. To this end, we propose TangoGS, which introduces a novel pixel-capture-based learning budget mechanism to derive the initial model size and dynamically controls Gaussian density via training quality feedback, achieving compact cross-scale reconstruction without manual hyperparameter tuning. Experimental results demonstrate that TangoGS matches state-of-the-art PSNR on standard benchmarks while reducing the number of Gaussians by 48%. Furthermore, on large-scale scenes, it outperforms the second-best method by 0.54 dB in PSNR, delivering superior rendering efficiency alongside enhanced reconstruction quality.
📝 Abstract
3D Gaussian Splatting reconstructs a scene as a collection of Gaussian primitives from a set of posed photographs called the capture. The number of primitives used to represent the scene affects reconstruction quality, storage, and rendering cost. How to select this number automatically across capture scales remains unresolved: configurations effective on standard benchmarks can leave larger captures with too few Gaussians to reconstruct fine details. We observe that the surface to represent, given by the capture's extent and resolution, is known before training, whereas its content complexity becomes apparent during training, through the reconstruction quality on the training views. We introduce TangoGS, which combines capture-derived model sizing with training-based adaptation: the capture determines the scale of the model, and training feedback determines its final size within that scale. Before training, TangoGS derives a learning allowance for model growth from the capture's total pixels after discounting views that re-observe the same scene points. During training, reconstruction quality guides how many Gaussians to add and remove. On 13 standard benchmark scenes, TangoGS matches the mean PSNR of the best-performing evaluated baseline, LeGS, with $48\%$ fewer Gaussians. On eight large captures, the same configuration automatically scales to larger models when necessary, achieving the highest mean PSNR among evaluated methods: $0.54$ dB above the runner-up with $2.3\times$ as many Gaussians. Together, capture-derived learning allowances and training-quality guided density control enable a state-of-the-art quality--size compromise across scene scales without retuning.