PIT-QMM: A Large Multimodal Model For No-Reference Point Cloud Quality Assessment

📅 2025-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the open problem of reference-free point cloud perceptual quality assessment by proposing the first end-to-end, large-scale multimodal evaluation framework. Methodologically, it fuses textual descriptions, 2D projection images, and raw 3D point clouds into a cross-modal joint encoder, leveraging attention mechanisms for deep alignment and complementary modeling across modalities—enabling both quality score prediction and distortion region localization. Key contributions include: (1) the first adaptation of vision-language model paradigms to point cloud quality assessment; (2) interpretable, fine-grained distortion type classification coupled with spatial localization; and (3) consistent, significant improvements over state-of-the-art methods on major benchmarks, with higher training efficiency. Extensive experiments demonstrate the framework’s superior accuracy, robustness, and practical utility for human-in-the-loop quality analysis.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workSystems and Infrastructure for Web, Mobile and WoT: Web performance, measurement, and characterization
📝 Abstract
Large Multimodal Models (LMMs) have recently enabled considerable advances in the realm of image and video quality assessment, but this progress has yet to be fully explored in the domain of 3D assets. We are interested in using these models to conduct No-Reference Point Cloud Quality Assessment (NR-PCQA), where the aim is to automatically evaluate the perceptual quality of a point cloud in absence of a reference. We begin with the observation that different modalities of data - text descriptions, 2D projections, and 3D point cloud views - provide complementary information about point cloud quality. We then construct PIT-QMM, a novel LMM for NR-PCQA that is capable of consuming text, images and point clouds end-to-end to predict quality scores. Extensive experimentation shows that our proposed method outperforms the state-of-the-art by significant margins on popular benchmarks with fewer training iterations. We also demonstrate that our framework enables distortion localization and identification, which paves a new way forward for model explainability and interactivity. Code and datasets are available at https://www.github.com/shngt/pit-qmm.
Problem

Research questions and friction points this paper is trying to address.

Assessing 3D point cloud quality without reference models
Integrating text, image, and point cloud data for evaluation
Improving model explainability through distortion localization and identification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal model combining text, images, point clouds
End-to-end quality prediction without reference data
Enables distortion localization and identification capabilities
💼 Related Jobs
No related jobs found.
S
Shashank Gupta
The University of Texas at Austin
Gregoire Phillips
Gregoire Phillips
Senior Researcher, Ericsson Research
PrivacyCyber-Physical SystemsDistributed SystemsExtended RealityDigital Media
A
Alan C. Bovik
The University of Texas at Austin