Benchmarking and Learning Multi-Dimensional Quality Evaluator for Text-to-3D Generation

📅 2024-12-15
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 1
📄 PDF
🤖 AI Summary
Current text-to-3D generation evaluation suffers from two key limitations: (1) existing benchmarks lack fine-grained coverage across prompt categories and evaluation dimensions; and (2) metrics focus predominantly on single-view alignment (e.g., text–3D similarity), hindering holistic, multi-dimensional quality assessment. To address these, we introduce MATE-3D—the first fine-grained, multi-dimensional evaluation benchmark—featuring 8 prompt categories, 1,280 textured meshes, and 107,520 human annotations. We further propose HyperScore, a learnable multi-dimensional evaluator that employs a hypernetwork to dynamically generate dimension-specific mapping functions for geometry, texture, semantics, layout, and more. This enables the first end-to-end, collaborative multi-dimensional evaluation of text-to-3D generation and establishes a novel paradigm of dimension-adaptive assessment. Experiments show HyperScore achieves a 32.7% improvement in correlation with human judgments over CLIPScore on MATE-3D. Both code and data are publicly released, establishing MATE-3D and HyperScore as community-standard evaluation tools.

Technology Category

Computer Vision: 3D Computer VisionNatural Language Processing: GenerationMachine Learning: Large Multimodal Models (LMMs)

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workUser Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating success
📝 Abstract
Text-to-3D generation has achieved remarkable progress in recent years, yet evaluating these methods remains challenging for two reasons: i) Existing benchmarks lack fine-grained evaluation on different prompt categories and evaluation dimensions. ii) Previous evaluation metrics only focus on a single aspect (e.g., text-3D alignment) and fail to perform multi-dimensional quality assessment. To address these problems, we first propose a comprehensive benchmark named MATE-3D. The benchmark contains eight well-designed prompt categories that cover single and multiple object generation, resulting in 1,280 generated textured meshes. We have conducted a large-scale subjective experiment from four different evaluation dimensions and collected 107,520 annotations, followed by detailed analyses of the results. Based on MATE-3D, we propose a novel quality evaluator named HyperScore. Utilizing hypernetwork to generate specified mapping functions for each evaluation dimension, our metric can effectively perform multi-dimensional quality assessment. HyperScore presents superior performance over existing metrics on MATE-3D, making it a promising metric for assessing and improving text-to-3D generation. The project is available at https://mate-3d.github.io/.
Problem

Research questions and friction points this paper is trying to address.

Lack of fine-grained evaluation benchmarks for text-to-3D generation
Existing metrics fail to assess multi-dimensional quality aspects
Need for a comprehensive evaluator to improve text-to-3D assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Comprehensive benchmark MATE-3D for fine-grained evaluation
HyperScore evaluator with hypernetwork for multi-dimensional assessment
Large-scale subjective experiment with 107,520 annotations
🔎 Similar Papers
No similar papers found.
Shanghai Jiao Tong University | University of Missouri-Kansas City
Yujie Zhang
Yujie Zhang
Shanghai Jiao tong University
3D Quality AssessmentGeometry Processing3D Reconstruction
B
Bingyang Cui
Shanghai Jiao Tong University
Q
Qi Yang
University of Missouri-Kansas City
Z
Zhu Li
University of Missouri-Kansas City
Yiling Xu
Yiling Xu
Shanghai Jiaotong University