Score
Designs and builds annotation schemas, tooling, and datasets that align 3D geometry (meshes, point clouds, or reconstructed scenes) and visual views with semantic labels and natural-language descriptions. Uses those 3D semantic and vision–language annotations to train and evaluate systems that infer object semantics, identify meaningful interaction regions, predict affordances and interaction constraints, and supply 3D-grounded cues and priors for downstream models.
This paper addresses two core challenges: the difficulty of transferring semantic information from 2D images to 3D point clouds, and the disconnection between implicit guidance and explicit modeling. It is the first work to systematically characterize semantics in point clouds as serving dual roles—implicit decision guidance and explicit structural modeling—and establishes a cross-task semantic fusion taxonomy spanning autonomous driving, surveying, and architecture. The authors propose a unified analytical framework integrating multimodal alignment, 3D semantic segmentation, scene graph reasoning, and cross-domain transfer learning, validated through comparative experiments on mainstream public benchmarks. Contributions include: (1) a structured literature repository; (2) a continuously updated GitHub resource platform; and (3) a standardized review benchmark and practical guideline for semantic point cloud research—providing both theoretical foundations and technical references for future method development and real-world deployment.
Existing 3D assets often lack semantic and physical interaction knowledge, hindering robots’ ability to understand where and how to manipulate objects. This work proposes AnnotateAnything, a framework that for the first time enables large-scale, fully automatic, and unified generation of multi-type robotic manipulation annotations. By integrating vision-language reasoning with physical constraints, the method combines geometric optimization, executable trajectory generation, and large-scale parallel physics simulation to efficiently produce structured, diverse, and physically feasible manipulation annotations. Experiments demonstrate that the framework substantially improves annotation efficiency and task success rates, while effectively supporting downstream applications such as functional region detection, robotic visual question answering, and vision-based instruction fine-tuning.
To address the high cost and low accuracy of manual annotation in 3D scene understanding, this paper proposes SCANnotate++, the first high-fidelity synthetic annotation paradigm leveraging automated CAD model retrieval and 9D pose estimation. SCANnotate++ precisely matches ScanNet++ v1 scenes against a large-scale CAD library and refines alignments via point cloud completion (PCN) and iterative closest point (ICP) optimization, yielding accurate, instance-level 3D annotations. Experiments demonstrate that models trained on SCANnotate++ annotations achieve a 3.2% lower Chamfer distance on point cloud completion compared to human-annotated baselines, and attain a 5.7% higher recall in single-view CAD retrieval and alignment. This work provides the first empirical validation that synthetic annotations can surpass manual ones in downstream performance, significantly improving model generalization. All annotations, code, and models are publicly released to advance cost-effective, high-accuracy 3D supervised learning.
Cross-dataset inconsistency in part definitions and lack of semantic naming hinder generalization in 3D part segmentation. Method: We propose the first end-to-end semantic naming segmentation framework, introducing *partlets*—learnable implicit part representations—and a geometric-visual-linguistic multimodal alignment mechanism. Our approach jointly optimizes implicit part fields, multi-view features, and LLM-generated functional descriptions, achieving fine-grained alignment via bipartite matching and text-geometry joint embedding. Contributions/Results: (1) The first unified part ontology covering PartNet, 3DCoMPaT++, and Find3D (1,794 categories); (2) Zero-shot semantic naming with calibrated confidence scoring; (3) Two dedicated evaluation metrics and the new Tex-Parts benchmark. Our method achieves state-of-the-art performance on PartNet and other major benchmarks.
This work addresses open-vocabulary semantic segmentation in 3D scenes, achieving the first end-to-end open-vocabulary segmentation over the complete 3D volumetric space of both NeRF and 3D Gaussian Splatting (3DGS), overcoming the limitation of prior methods that produce only 2D masks. The proposed method introduces point-level language embedding field supervision, a cross-representation semantic transfer mechanism (NeRF → 3DGS), and the first geometry-semantic joint 3D query evaluation protocol. It integrates 3D point cloud language embedding learning, CLIP feature distillation, and voxel-level semantic querying. Evaluated on ScanNet and Objaverse, it achieves state-of-the-art 3D semantic segmentation accuracy while enabling real-time rendering (>60 FPS). This work establishes a novel paradigm for open-vocabulary 3D understanding.
This work addresses language-query-based embodied vision-language navigation under open-set conditions. Methodologically, it introduces a 3D Open-Set Semantic Instance Mapping (O3D-SIM) framework that integrates multimodal foundation models—specifically CLIP for cross-modal alignment and zero-shot object recognition, and SAM for image-level segmentation—combined with SLAM-based pose estimation and point-cloud instance clustering to construct an open-set, instance-level, semantically enriched 3D map. The core contribution is the first realization of open-set 3D instance semantic mapping capable of supporting language queries involving previously unseen object categories, thereby overcoming the limitations of conventional closed-set semantic mapping. Experimental results demonstrate substantial improvements in language-guided navigation success rates; qualitative analysis further confirms strong generalization capability and interpretability of the generated maps.
Current visual language model (VLM) training is hindered by the scarcity of high-quality annotated data that jointly captures unified spatial coordinates, open-vocabulary semantics, structural attributes, and topological relationships. Moreover, conventional annotation tools suffer from limited expressiveness, a disconnect between annotation and training pipelines, and poor reusability. To address these challenges, this work proposes ScreenAnnotator, which introduces a unified atomic annotation schema and integrates an online policy-based annotation loop with an embedded Bayesian verifier alongside a template-driven multitask data synthesis mechanism. This framework enables efficient and reusable construction of visual reasoning datasets. Evaluated on flowchart and GUI screenshot annotations, the system achieves acceptance rates of 99.7% and 77%, respectively, with consistently decreasing per-image annotation time. Fine-tuning VLMs on the generated data yields a 76.1% accuracy on flowchart understanding tasks, representing an absolute improvement of 35.1 percentage points.
This work addresses the challenge of aligning 3D object functional regions with natural language instructions under open-vocabulary settings by proposing a two-stage cross-modal framework. It first leverages a large language model to generate part-aware completion instructions that enrich semantic representations. Subsequently, it jointly optimizes cross-object geometric consistency and intra-object semantic alignment through Adaptive Prototype Aggregation (APA) and Intra-Object Relation Modeling (IORM). Evaluated on a newly constructed benchmark as well as two existing datasets, the proposed method significantly outperforms current state-of-the-art approaches, demonstrating strong effectiveness and generalization capability in open-vocabulary 3D functional grounding tasks.
Existing 3D vision-language models suffer from severe degradation of geometric information in their intermediate representations due to the scarcity of paired 3D-text data and reliance solely on language token supervision. To address this, this work proposes a lightweight, feature-level alignment regularization method that explicitly enforces consistency between intermediate point cloud tokens and the original visual input during language modeling via a consistency loss, thereby preserving fine-grained geometric semantics. This approach introduces, for the first time in 3D vision-language models, a feature-level alignment mechanism that requires training only an alignment projector and LoRA adapters, ensuring high efficiency with minimal computational overhead. Experiments demonstrate consistent improvements: average classification accuracy increases by 2.08% on ModelNet40 and Objaverse, open-vocabulary classification performance improves by 7.50%, and 3D captioning quality rises by 4.88%.
This work addresses the challenge of accurately localizing geometric entities—such as edges and faces—on 3D meshes using natural language, a task hindered by view sensitivity, occlusion, and perspective distortion. The authors propose MV-GEL, the first framework to achieve high-precision language-driven localization using only raw mesh inputs. MV-GEL introduces GELviews, a module that ranks multi-view observability based on linguistic cues to select the optimal rendering viewpoint. It then leverages a vision-language model for 2D segmentation and employs geometry-aware ray casting to map the results back onto the original 3D mesh. Evaluated on standard benchmarks, MV-GEL improves face-level IoU by 1.7× and edge-level F1 score by over 4.5×, substantially outperforming CLIP-based baselines and random view sampling strategies.