Score
Designs and builds annotation tools, workflows, and labeled datasets for 3D scenes (point clouds and meshes), supporting point-level labels, polygon traces on image projections, and alignment of annotations across agents and modalities. Implements distortion-aware labeling and quality-assessment artifacts—such as point-cloud distortion taxonomies, discrete distortion-severity categories, structured natural-language descriptions, and QA procedures—to produce pixel-accurate, consistent annotations and manage annotation quality.
This work addresses the limited interpretability of existing point cloud quality assessment (PCQA) methods, which typically predict only overall subjective scores without elucidating the underlying perceptual degradations. To bridge this gap, we introduce a novel dataset featuring the first distortion taxonomy specifically designed for point clouds, accompanied by multi-level distortion severity labels, discrete quality categories, and structured natural language descriptions. Leveraging this dataset, we conduct zero-shot and fine-tuned experiments using multimodal models augmented with human-perception-aligned natural language generation. Our results demonstrate that incorporating distortion-aware supervision significantly enhances the lexical and semantic alignment between generated descriptions and human annotations, thereby advancing interpretable PCQA.
Existing point cloud quality assessment methods typically predict only scalar scores, which are insufficient to support diagnostic needs such as defect identification, attribution, and interpretability. To address this limitation, this work proposes PointQ-Bench—the first benchmark specifically designed for diagnostic and interpretable point cloud quality evaluation—comprising 3,083 samples with multidimensional annotations. The benchmark introduces the SSFRQ-5D five-dimensional evaluation protocol, enabling four key tasks: anomaly-aware assessment, defect diagnosis, usability grading, and open-ended quality reporting. Leveraging real-world scans, simulated distortions, and AI-generated data, the study systematically evaluates 14 models, including both vision-language foundation models and traditional approaches. Experimental results reveal a significant performance gap between coarse-grained perception and fine-grained diagnosis, and demonstrate that powerful 2D multimodal large models consistently outperform specialized 3D architectures.
To address the high cost and low accuracy of manual annotation in 3D scene understanding, this paper proposes SCANnotate++, the first high-fidelity synthetic annotation paradigm leveraging automated CAD model retrieval and 9D pose estimation. SCANnotate++ precisely matches ScanNet++ v1 scenes against a large-scale CAD library and refines alignments via point cloud completion (PCN) and iterative closest point (ICP) optimization, yielding accurate, instance-level 3D annotations. Experiments demonstrate that models trained on SCANnotate++ annotations achieve a 3.2% lower Chamfer distance on point cloud completion compared to human-annotated baselines, and attain a 5.7% higher recall in single-view CAD retrieval and alignment. This work provides the first empirical validation that synthetic annotations can surpass manual ones in downstream performance, significantly improving model generalization. All annotations, code, and models are publicly released to advance cost-effective, high-accuracy 3D supervised learning.
Scientific image annotation projects face cross-domain managerial challenges—including scarce data acquisition, inefficient resource allocation, inadequate annotator training, and pronounced human bias. To address these issues, this paper proposes the first general-purpose framework for preparing scientific image annotation projects. The framework systematically integrates objective definition, data availability assessment, multi-role team configuration, bias mitigation strategies, and an iterative annotator training mechanism, complemented by a recommended toolchain supporting integrated project management, quality control, and collaborative annotation. A novel closed-loop workflow—comprising bias detection, feedback integration, and retraining—is introduced to significantly enhance annotation consistency and efficiency. Empirical evaluation across multiple disciplines demonstrates that the framework reduces annotation costs by over 20%, improves project success rates, and strengthens knowledge base construction quality—thereby filling a critical research gap in standardized preparation guidelines for complex scientific image annotation.
Current text-to-3D generation evaluation suffers from two key limitations: (1) existing benchmarks lack fine-grained coverage across prompt categories and evaluation dimensions; and (2) metrics focus predominantly on single-view alignment (e.g., text–3D similarity), hindering holistic, multi-dimensional quality assessment. To address these, we introduce MATE-3D—the first fine-grained, multi-dimensional evaluation benchmark—featuring 8 prompt categories, 1,280 textured meshes, and 107,520 human annotations. We further propose HyperScore, a learnable multi-dimensional evaluator that employs a hypernetwork to dynamically generate dimension-specific mapping functions for geometry, texture, semantics, layout, and more. This enables the first end-to-end, collaborative multi-dimensional evaluation of text-to-3D generation and establishes a novel paradigm of dimension-adaptive assessment. Experiments show HyperScore achieves a 32.7% improvement in correlation with human judgments over CLIPScore on MATE-3D. Both code and data are publicly released, establishing MATE-3D and HyperScore as community-standard evaluation tools.
This work addresses the open problem of reference-free point cloud perceptual quality assessment by proposing the first end-to-end, large-scale multimodal evaluation framework. Methodologically, it fuses textual descriptions, 2D projection images, and raw 3D point clouds into a cross-modal joint encoder, leveraging attention mechanisms for deep alignment and complementary modeling across modalities—enabling both quality score prediction and distortion region localization. Key contributions include: (1) the first adaptation of vision-language model paradigms to point cloud quality assessment; (2) interpretable, fine-grained distortion type classification coupled with spatial localization; and (3) consistent, significant improvements over state-of-the-art methods on major benchmarks, with higher training efficiency. Extensive experiments demonstrate the framework’s superior accuracy, robustness, and practical utility for human-in-the-loop quality analysis.
This paper addresses text-driven scene-consistent image generation: synthesizing images that simultaneously preserve geometric/appearance fidelity to a reference scene graph and accurately realize textual descriptions of target entities and their spatial relationships. To overcome the trade-off between these competing objectives in existing methods, we propose a geometry-guided diffusion framework comprising: (1) a multi-view geometric modeling pipeline for constructing scene-consistent training data; (2) self-supervised spatial regularization incorporating cross-view geometric constraints; and (3) a scene-text joint attention mechanism. Notably, this work is the first to explicitly integrate geometric priors into attention optimization for text-to-image generation. On our newly established benchmark, our method achieves +12.6% CLIP-Scene score, +9.3% TIFA score, and 78.4% human preference rate, demonstrating strong capability in generating complex geometric compositions.
Existing point cloud generation evaluation metrics—such as Chamfer Distance—are highly sensitive to geometric imperfections and lack robustness, failing to accurately quantify both local shape consistency and global fidelity. To address these limitations, this work proposes: (1) two novel evaluation metrics—Density-Aware Chamfer Distance (DCD) and Surface Normal Consistency (SNC)—designed to better discriminate sampling non-uniformity and normal vector distortion; and (2) Diffusion Point Transformer, a diffusion-based generative architecture leveraging serialized patch-wise attention, augmented with sample alignment preprocessing to enhance local structural modeling. Evaluated on ShapeNet, our method achieves state-of-the-art generation quality, significantly outperforming leading baselines across multiple metrics. The implementation is publicly available.
This work addresses the performance bottlenecks in 3D object annotation caused by occlusion, viewpoint variation, and spatial complexity by proposing Tri-MARF, a novel framework that introduces, for the first time, a trimodal multi-agent collaboration mechanism to jointly process 2D multi-view images, 3D point clouds, and textual descriptions. By decoupling and co-optimizing visual-language understanding, information selection, and semantic-geometric alignment capabilities, the method significantly enhances both annotation accuracy and scalability. Evaluated on the Objaverse and LVIS datasets, the model achieves a CLIPScore of 88.7 and ViLT R@5 accuracies of 45.2 and 43.8, respectively, while demonstrating high-throughput performance by annotating 12,000 objects per hour on a single A100 GPU.