🤖 AI Summary
Current text-to-3D generation evaluation suffers from two key limitations: (1) existing benchmarks lack fine-grained coverage across prompt categories and evaluation dimensions; and (2) metrics focus predominantly on single-view alignment (e.g., text–3D similarity), hindering holistic, multi-dimensional quality assessment. To address these, we introduce MATE-3D—the first fine-grained, multi-dimensional evaluation benchmark—featuring 8 prompt categories, 1,280 textured meshes, and 107,520 human annotations. We further propose HyperScore, a learnable multi-dimensional evaluator that employs a hypernetwork to dynamically generate dimension-specific mapping functions for geometry, texture, semantics, layout, and more. This enables the first end-to-end, collaborative multi-dimensional evaluation of text-to-3D generation and establishes a novel paradigm of dimension-adaptive assessment. Experiments show HyperScore achieves a 32.7% improvement in correlation with human judgments over CLIPScore on MATE-3D. Both code and data are publicly released, establishing MATE-3D and HyperScore as community-standard evaluation tools.
📝 Abstract
Text-to-3D generation has achieved remarkable progress in recent years, yet evaluating these methods remains challenging for two reasons: i) Existing benchmarks lack fine-grained evaluation on different prompt categories and evaluation dimensions. ii) Previous evaluation metrics only focus on a single aspect (e.g., text-3D alignment) and fail to perform multi-dimensional quality assessment. To address these problems, we first propose a comprehensive benchmark named MATE-3D. The benchmark contains eight well-designed prompt categories that cover single and multiple object generation, resulting in 1,280 generated textured meshes. We have conducted a large-scale subjective experiment from four different evaluation dimensions and collected 107,520 annotations, followed by detailed analyses of the results. Based on MATE-3D, we propose a novel quality evaluator named HyperScore. Utilizing hypernetwork to generate specified mapping functions for each evaluation dimension, our metric can effectively perform multi-dimensional quality assessment. HyperScore presents superior performance over existing metrics on MATE-3D, making it a promising metric for assessing and improving text-to-3D generation. The project is available at https://mate-3d.github.io/.