SpatialQ: Understanding 3D Gaussian Splatting Scene Quality via Visual-based MLLM

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing image quality assessment (IQA) methods struggle to effectively evaluate the quality of 3D Gaussian Splatting (3DGS) scenes due to their neglect of three-dimensional structure and cross-view consistency. This work proposes the first multimodal quality assessment framework tailored for 3DGS, integrating rendered images, depth maps, point clouds, and camera parameters. The framework employs a VGGT-enhanced encoder, multi-view feature aggregation, and joint geometric modeling to extract structure-aware representations, and further introduces a grounded multimodal large language model based on Qwen for quality regression and reasoning. By explicitly incorporating geometric structure and multi-view consistency into 3DGS quality evaluation, the proposed method substantially outperforms existing 2D IQA approaches and general-purpose multimodal large language models, enabling reliable assessment of both perceptual fidelity and 3D structural integrity.
📝 Abstract
3D Gaussian Splatting (3DGS) has emerged as an effective representation for novel view synthesis and 3D scene reconstruction, creating an increasing demand for reliable quality assessment. Unlike conventional image quality assessment (IQA), the quality of a 3DGS scene depends not only on the perceptual fidelity of rendered views, but also on scene-level factors such as spatial structure and cross-view consistency. Existing IQA methods are limited by their reliance on 2D perceptual cues, whereas general multimodal large language models (MLLMs) are not designed for stable quality regression and may produce unreliable judgments. To address these limitations, a multimodal quality assessment framework is developed for 3DGS scene understanding. First, a 3D-aware quality representation learning framework is introduced by augmenting a VGGT-based encoder with a dedicated quality head. Multi-view images are encoded into view-specific features and aggregated to capture cross-view consistency, while geometric cues are incorporated through joint modeling of depth and point-cloud-related structural information, enabling the learning of structure-aware quality representations beyond appearance-driven features. Second, a grounded multimodal reasoning mechanism is constructed by jointly feeding original images, depth maps, point cloud renderings, and camera parameters into a Qwen-based MLLM.
Problem

Research questions and friction points this paper is trying to address.

3D Gaussian Splatting
quality assessment
spatial structure
cross-view consistency
multimodal reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D Gaussian Splatting
quality assessment
multimodal large language model
cross-view consistency
structure-aware representation
🔎 Similar Papers
No similar papers found.