vision-language 3d annotation

Designs and builds annotation schemas, tooling, and datasets that align 3D geometry (meshes, point clouds, or reconstructed scenes) and visual views with semantic labels and natural-language descriptions. Uses those 3D semantic and vision–language annotations to train and evaluate systems that infer object semantics, identify meaningful interaction regions, predict affordances and interaction constraints, and supply 3D-grounded cues and priors for downstream models.

vision-language3dannotation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing 3D assets often lack semantic and physical interaction knowledge, hindering robots’ ability to understand where and how to manipulate objects. This work proposes AnnotateAnything, a framework that for the first time enables large-scale, fully automatic, and unified generation of multi-type robotic manipulation annotations. By integrating vision-language reasoning with physical constraints, the method combines geometric optimization, executable trajectory generation, and large-scale parallel physics simulation to efficiently produce structured, diverse, and physically feasible manipulation annotations. Experiments demonstrate that the framework substantially improves annotation efficiency and task success rates, while effectively supporting downstream applications such as functional region detection, robotic visual question answering, and vision-based instruction fine-tuning.

3D asset annotationinteractive labelingphysical affordances

Leveraging Automatic CAD Annotations for Supervised Learning in 3D Scene Understanding

Apr 18, 2025
YR
Yuchen Rao
🏛️ Graz Univ. of Technology | Ecole des Ponts et Chaussees | IP Paris | CNRS

To address the high cost and low accuracy of manual annotation in 3D scene understanding, this paper proposes SCANnotate++, the first high-fidelity synthetic annotation paradigm leveraging automated CAD model retrieval and 9D pose estimation. SCANnotate++ precisely matches ScanNet++ v1 scenes against a large-scale CAD library and refines alignments via point cloud completion (PCN) and iterative closest point (ICP) optimization, yielding accurate, instance-level 3D annotations. Experiments demonstrate that models trained on SCANnotate++ annotations achieve a 3.2% lower Chamfer distance on point cloud completion compared to human-annotated baselines, and attain a 5.7% higher recall in single-view CAD retrieval and alignment. This work provides the first empirical validation that synthetic annotations can surpass manual ones in downstream performance, significantly improving model generalization. All annotations, code, and models are publicly released to advance cost-effective, high-accuracy 3D supervised learning.

Generating accurate 3D annotations for deep learning modelsReducing annotation costs while improving model performanceUsing synthetic CAD models as ground truth for training

Name That Part: 3D Part Segmentation and Naming

Dec 19, 2025
SP
Soumava Paul
🏛️ Johns Hopkins University

Cross-dataset inconsistency in part definitions and lack of semantic naming hinder generalization in 3D part segmentation. Method: We propose the first end-to-end semantic naming segmentation framework, introducing *partlets*—learnable implicit part representations—and a geometric-visual-linguistic multimodal alignment mechanism. Our approach jointly optimizes implicit part fields, multi-view features, and LLM-generated functional descriptions, achieving fine-grained alignment via bipartite matching and text-geometry joint embedding. Contributions/Results: (1) The first unified part ontology covering PartNet, 3DCoMPaT++, and Find3D (1,794 categories); (2) Zero-shot semantic naming with calibrated confidence scoring; (3) Two dedicated evaluation metrics and the new Tex-Parts benchmark. Our method achieves state-of-the-art performance on PartNet and other major benchmarks.

Aligns inconsistent 3D part annotations across datasetsCreates a unified part ontology for open-vocabulary matchingSegments and names 3D parts via text and geometry alignment

Rethinking Open-Vocabulary Segmentation of Radiance Fields in 3D Space

Aug 14, 2024
HL
Hyunjee Lee
🏛️ Yonsei University

This work addresses open-vocabulary semantic segmentation in 3D scenes, achieving the first end-to-end open-vocabulary segmentation over the complete 3D volumetric space of both NeRF and 3D Gaussian Splatting (3DGS), overcoming the limitation of prior methods that produce only 2D masks. The proposed method introduces point-level language embedding field supervision, a cross-representation semantic transfer mechanism (NeRF → 3DGS), and the first geometry-semantic joint 3D query evaluation protocol. It integrates 3D point cloud language embedding learning, CLIP feature distillation, and voxel-level semantic querying. Evaluated on ScanNet and Objaverse, it achieves state-of-the-art 3D semantic segmentation accuracy while enabling real-time rendering (>60 FPS). This work establishes a novel paradigm for open-vocabulary 3D understanding.

Develop new 3D querying and evaluation methodsEnable real-time rendering with 3D Gaussian splattingImprove 3D semantic segmentation in radiance fields

Open-set 3D semantic instance maps for vision language navigation – O3D-SIM

Apr 27, 2024
LN
Laksh Nanwani
🏛️ International Institute of Information Technology, Hyderabad | Hasan Kalyoncu University

This work addresses language-query-based embodied vision-language navigation under open-set conditions. Methodologically, it introduces a 3D Open-Set Semantic Instance Mapping (O3D-SIM) framework that integrates multimodal foundation models—specifically CLIP for cross-modal alignment and zero-shot object recognition, and SAM for image-level segmentation—combined with SLAM-based pose estimation and point-cloud instance clustering to construct an open-set, instance-level, semantically enriched 3D map. The core contribution is the first realization of open-set 3D instance semantic mapping capable of supporting language queries involving previously unseen object categories, thereby overcoming the limitations of conventional closed-set semantic mapping. Experimental results demonstrate substantial improvements in language-guided navigation success rates; qualitative analysis further confirms strong generalization capability and interpretability of the generated maps.

Creating 3D semantic instance maps for language-guided navigation tasksEnhancing success rates of vision-language navigation through instance embeddingsImproving object recognition robustness using open-set foundational models

Latest Papers

What's happening recently
View more

Current visual language model (VLM) training is hindered by the scarcity of high-quality annotated data that jointly captures unified spatial coordinates, open-vocabulary semantics, structural attributes, and topological relationships. Moreover, conventional annotation tools suffer from limited expressiveness, a disconnect between annotation and training pipelines, and poor reusability. To address these challenges, this work proposes ScreenAnnotator, which introduces a unified atomic annotation schema and integrates an online policy-based annotation loop with an embedded Bayesian verifier alongside a template-driven multitask data synthesis mechanism. This framework enables efficient and reusable construction of visual reasoning datasets. Evaluated on flowchart and GUI screenshot annotations, the system achieves acceptance rates of 99.7% and 77%, respectively, with consistently decreasing per-image annotation time. Fine-tuning VLMs on the generated data yields a 76.1% accuracy on flowchart understanding tasks, representing an absolute improvement of 35.1 percentage points.

annotation-tool bottleneckdata annotationgrounded representation

This work addresses the challenge of aligning 3D object functional regions with natural language instructions under open-vocabulary settings by proposing a two-stage cross-modal framework. It first leverages a large language model to generate part-aware completion instructions that enrich semantic representations. Subsequently, it jointly optimizes cross-object geometric consistency and intra-object semantic alignment through Adaptive Prototype Aggregation (APA) and Intra-Object Relation Modeling (IORM). Evaluated on a newly constructed benchmark as well as two existing datasets, the proposed method significantly outperforms current state-of-the-art approaches, demonstrating strong effectiveness and generalization capability in open-vocabulary 3D functional grounding tasks.

3D affordance groundinggeometric alignmentopen-vocabulary

Existing 3D vision-language models suffer from severe degradation of geometric information in their intermediate representations due to the scarcity of paired 3D-text data and reliance solely on language token supervision. To address this, this work proposes a lightweight, feature-level alignment regularization method that explicitly enforces consistency between intermediate point cloud tokens and the original visual input during language modeling via a consistency loss, thereby preserving fine-grained geometric semantics. This approach introduces, for the first time in 3D vision-language models, a feature-level alignment mechanism that requires training only an alignment projector and LoRA adapters, ensuring high efficiency with minimal computational overhead. Experiments demonstrate consistent improvements: average classification accuracy increases by 2.08% on ModelNet40 and Objaverse, open-vocabulary classification performance improves by 7.50%, and 3D captioning quality rises by 4.88%.

3D Vision-Language Modelsgeometric information degradationinefficient 3D data utilization

This work addresses the challenge of accurately localizing geometric entities—such as edges and faces—on 3D meshes using natural language, a task hindered by view sensitivity, occlusion, and perspective distortion. The authors propose MV-GEL, the first framework to achieve high-precision language-driven localization using only raw mesh inputs. MV-GEL introduces GELviews, a module that ranks multi-view observability based on linguistic cues to select the optimal rendering viewpoint. It then leverages a vision-language model for 2D segmentation and employs geometry-aware ray casting to map the results back onto the original 3D mesh. Evaluated on standard benchmarks, MV-GEL improves face-level IoU by 1.7× and edge-level F1 score by over 4.5×, substantially outperforming CLIP-based baselines and random view sampling strategies.

3D meshesgeometric entity localizationmulti-view perception

Hot Scholars

WW

Wenping Wang

Texas A&M University
Computer GraphicsGeometric Computing
EL

Ed Li

Yale University
agentic systemsai4scienceautoML
MB

Muyi Bao

Carneige Mellon University
ZW

Zimu Wang

Tsinghua University
recommendation
FR

Fabio Remondino

3D Optical Metrology - Bruno Kessler Foundation
photogrammetry3D modelingAI