🤖 AI Summary
This study addresses the prohibitive depth-sorting overhead and noise artifacts arising from stochastic transparency in 3D Gaussian Splatting. To this end, we propose a sort-free blended pointillism sampling framework that introduces an adaptive fusion of primitive- and fragment-level sampling strategies. By leveraging attribute information to guide the sampling process and incorporating a Gaussian-aware spatiotemporal reconstruction network, our method effectively suppresses rendering noise. As a core contribution, this framework achieves interactive-rate, temporally stable rendering on mobile devices without requiring retraining. Furthermore, under specific configurations, it yields visual quality superior to existing baseline methods.
📝 Abstract
Conventional 3D Gaussian Splatting (3DGS) requires depth sorting and ordered alpha blending to correctly render overlapping Gaussian primitives. Stochastic transparency enables sorting-free rendering by replacing fractional alpha contributions with discrete stochastic visibility samples, but produces substantial spatial and temporal noise at low sample counts. We refer to this conversion from continuous Gaussian splats to discrete visibility samples as \textit{Gaussian Stippling}. Based on this, we present an efficient order-independent rendering and reconstruction framework that operates directly on unmodified 3DGS assets. Our method adaptively integrates primitive-based and fragment-based stippling, leveraging their complementary strengths across different rendering regimes to significantly improve rendering throughput. To recover high-quality images from sparse stochastic samples, we further introduce a lightweight Gaussian-aware spatiotemporal reconstruction network. By exploiting the Gaussian attributes retained by each stipple, the network aggregates structured stochastic clues across both space and time, effectively suppressing stippling noise. Experiments show that our hybrid Gaussian stippling method, coupled with a spatiotemporal reconstruction network trained on diverse scenes, generalizes to unseen scenes and enables interactive, temporally stable, and visually plausible rendering on mobile devices without retraining or preprocessing. With scene-specific training and appropriately scaled sampling and network capacity, our method further outperforms the baselines in visual quality, offering a high-fidelity configuration for quality-prioritized applications.