Lang3DSeg: Annotation-Free Open-Vocabulary 3D Segmentation with Point Transformers

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of existing open-vocabulary outdoor LiDAR segmentation methods, including reliance on voxel convolutions, pre-training, and label noise introduced by 2D projections. To overcome these challenges, this work proposes the first annotation-free, zero-shot 3D segmentation framework based on a point Transformer architecture. Methodologically, it enables end-to-end training from scratch via knowledge distillation from vision-language models (VLMs), while introducing a class-prior mask and an explicit depth distribution correction algorithm to effectively mitigate projection errors. The proposed framework achieves 52.8% and 41.4% mIoU on the nuScenes and SemanticKITTI benchmarks, respectively. Notably, it supports real-time inference without requiring VLM execution during deployment, offering a highly efficient solution for open-vocabulary 3D scene understanding.
πŸ“ Abstract
Accurate 3D semantic perception is critical for safe autonomous navigation. However, supervised LiDAR segmentation remains tied to closed taxonomies and to the cost of point-wise manual annotation. Open-vocabulary methods avoid that cost by projecting the output of 2D vision-language models onto LiDAR and distilling it into a 3D network. These methods rely almost exclusively on voxel-based sparse convolutions, and point transformers have so far been limited to indoor environments, where 3D data is dense and bounded. We present Lang3DSeg, which establishes a point transformer as the backbone for annotation-free open-vocabulary segmentation of outdoor 3D LiDAR, and is trained from scratch without geometric pre-training. This training paradigm necessitates addressing the inherent noise in 2D-to-3D label projections; specifically, naive projection often suffers from depth ambiguity, where points behind an object are erroneously assigned its semantic label. We therefore composite masks using an explicit class-priority rule and truncate each projected instance at the first gap in its depth distribution, correcting the projection error directly rather than averaging it over registered sequences. Lang3DSeg achieves 52.8% mIoU on nuScenes validation and 41.4% on SemanticKITTI, the highest among published annotation-free methods on both benchmarks. Every 3D semantic segmentation is on a single LiDAR sweep, and inference operates in real-time without running vision-language models.
Problem

Research questions and friction points this paper is trying to address.

3D semantic segmentation
open-vocabulary
LiDAR
point transformer
depth ambiguity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Open-Vocabulary 3D Segmentation
Point Transformers
Annotation-Free Learning
Depth Ambiguity Correction
Knowledge Distillation
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
C
Cigdem Kokenoz
Department of Automotive Engineering, Clemson University, Greenville, SC, USA
A
Amir Salarpour
School of Computing, Clemson University, Clemson, SC, USA
A
Alkim Domeke
School of Computing, Clemson University, Clemson, SC, USA
C
Christopher Salas
School of Computing, Clemson University, Clemson, SC, USA
Pedram MohajerAnsari
Pedram MohajerAnsari
Clemson University
Machine LearningOptimizationAutomotive Security
Long Cheng
Long Cheng
Associate Professor, School of Computing, Clemson University
Cyber SecurityInternet of ThingsWireless NetworksMobile Computing
Mert D. PesΓ©
Mert D. PesΓ©
Assistant Professor, Clemson University
Automotive SecurityAutomotive Privacy
B
Bing Li
Department of Automotive Engineering, Clemson University, Greenville, SC, USA