🤖 AI Summary
This work systematically investigates the semantic representation capacity of 3D Gaussian splatting for scene understanding. For this emerging representation, it presents the first comprehensive evaluation of geometric deep learning architectures—including point cloud networks and graph neural networks—on scene classification tasks, employing a combination of end-to-end training, linear probing, and clustering analysis. The study reveals significant performance disparities across model families when applied to Gaussian splatting data and demonstrates that geometry- and appearance-specific attributes inherent to Gaussians—such as covariance and opacity—substantially enhance representation quality. These findings establish a foundational basis for future semantic understanding tasks leveraging Gaussian splatting representations.
📝 Abstract
3D Gaussian Splatting (3DGS) is a recent approach for scene rendering. Although primarily designed for view synthesis, its potential for scene understanding tasks remains underexplored. In this work, we conduct a comparative evaluation of various geometric deep learning architectures for the classification of 3D scenes represented using Gaussian Splatting. We benchmark point-based and graph-based models across both traditional point cloud datasets and dedicated Gaussian Splatting datasets. Scenes are embedded into latent representations, which are evaluated through end-to-end classification, linear probing, and clustering analysis. Our study provides insight into the suitability of different geometry-aware architectures and input feature configurations for learning effective 3D Gaussian Splat representations. The results highlight consistent differences between architectural families and reveal the impact of Gaussian-specific attributes on the quality of representation.