Score
Designs and implements algorithms and processing pipelines that ingest, preprocess, and manipulate 3D point cloud data (e.g., LiDAR or depth-derived points) and related imagery to produce geometric representations such as meshes, Delaunay triangulations, normals, and compact point-based descriptors. Builds and evaluates methods for point cloud registration and alignment, camera–point-cloud projection and sensor calibration (including LiDAR reflectivity calibration), conversion between depth/images and point clouds, and semantic/instance segmentation and other representations for downstream 3D analysis.
Conventional geometric accuracy monitoring of unstructured point cloud data (PCD) suffers from error accumulation, high computational cost, and artifact generation due to reliance on registration and mesh reconstruction. Method: This paper proposes a preprocessing-free, intrinsic geometry-aware monitoring method that directly extracts intrinsic geometric features from raw PCD via Laplacian eigenmaps and geodesic distance computation. It incorporates an unsupervised threshold selection mechanism and a dual-strategy feature learning framework to enable robust defect identification. Contribution/Results: Experimental results demonstrate that the method achieves efficient and accurate detection of diverse geometric defects—including dents, bulges, and misalignments—without registration or reconstruction. It maintains high precision (>96% F1-score) and strong stability across heterogeneous manufacturing scenarios (e.g., additive manufacturing, CNC machining), reducing monitoring time by up to 78% compared to state-of-the-art approaches while eliminating reconstruction-induced artifacts. The framework significantly enhances both efficiency and reliability in industrial PCD-based quality inspection.
Current deep learning methods for 3D point clouds face significant deployment challenges in urban and environmental applications: existing research overemphasizes architectural innovation while neglecting practical constraints—including ultra-large scale, non-uniform density, multimodality, and scene diversity. This paper pioneers a reverse survey paradigm (“application scenario → algorithmic requirement”) to systematically evaluate the adaptability and limitations of representative architectures—PointNet, PointCNN, and graph neural networks—across key tasks: point cloud completion, registration, semantic segmentation, and reconstruction. We benchmark their performance on mainstream datasets, identifying insufficient generalization as a root cause. Furthermore, we explicitly characterize the gap between algorithmic innovation and engineering deployment. Finally, we propose a co-design pathway that jointly optimizes model lightweighting, cross-density robustness, and multimodal fusion—thereby providing both theoretical foundations and practical guidelines for deployable point cloud intelligence.
To address data scarcity, poor generalization, and high computational cost in few-shot LiDAR point cloud semantic segmentation, this paper proposes a LiDAR-only point-to-plane projection fusion framework. The method introduces a differentiable point-to-plane projection operator that maps 3D point clouds onto multiple geometry-preserving 2D representations; integrates geometry-aware data augmentation to mitigate class imbalance; and efficiently extracts complementary features in 2D before back-projecting them to 3D for semantic prediction. Crucially, it requires no RGB images or additional annotations. Evaluated on SemanticKITTI and PandaSet, our approach significantly outperforms existing few-shot methods—especially under extreme label scarcity (1%–10% annotated data)—while maintaining high efficiency and strong cross-scene generalization.
This work addresses the limited representational capacity of point cloud learning by proposing a cross-modal representation learning framework that incorporates 2D visual priors. Methodologically, it innovatively leverages pre-trained 2D vision models to guide 3D point cloud network training via feature alignment—enabling effective 2D→3D knowledge transfer without naïve modality conversion. The framework integrates a point cloud encoder-decoder architecture, self-supervised pre-training, and primitive-level supervised segmentation into a unified learning paradigm. Extensive experiments on ScanNet and S3DIS benchmarks demonstrate substantial improvements in point cloud segmentation and scene understanding performance, validating the efficacy of 2D semantic priors in enhancing 3D representation learning. The proposed approach establishes a scalable, multimodal representation learning pathway for efficient 3D perception in autonomous driving and robotics applications.
To address the high annotation cost of real 3D point cloud data—which severely limits model scalability and generalization in point cloud instance segmentation—this paper proposes a novel pretraining paradigm leveraging synthetic 3D data. We are the first to directly utilize the text-to-3D generative model Point-E to synthesize large-scale, photorealistic point cloud scenes with precise, instance-level annotations. Building upon this, we introduce a two-stage pretraining–fine-tuning framework tailored for downstream instance segmentation tasks (e.g., PointGroup). Our approach substantially reduces reliance on costly real-world annotations while significantly improving segmentation accuracy and cross-scene generalization on benchmarks such as ScanNet—outperforming both non-pretrained and image-level pretrained baselines. The core contribution lies in establishing a 3D generative model–driven, instance-aware synthetic data pretraining pipeline, offering a new paradigm for low-annotation-cost, robust 3D scene understanding.
This work addresses the inefficiency in loading and processing large-scale 3D point cloud data, which stems from heterogeneous formats and massive data volumes. To overcome these challenges, the authors propose .PcRecord, a unified binary storage format, coupled with a multi-stage parallel data processing pipeline that jointly optimizes storage compression and accelerated data loading. The proposed architecture is designed to support heterogeneous computing platforms, including GPU and Ascend accelerators. Evaluated across multiple benchmark datasets—ModelNet40, S3DIS, ShapeNet, KITTI, SUN RGB-D, and ScanNet—the approach achieves end-to-end speedups ranging from 2.23× to 25.4× compared to existing baseline methods, substantially enhancing the efficiency of point cloud processing workflows.
To address the high-cost reliance on LiDAR point clouds for digital twin modeling of critical infrastructure (e.g., hydraulic and energy systems), this paper proposes a low-cost RGB-D–driven method for automatic graph-structured model generation. The approach integrates stereo photogrammetry, deep learning–based instance segmentation, and user-customizable heuristic relation reasoning to achieve end-to-end mapping from raw images to structured graphs—where nodes represent physical devices and edges encode physical or functional connections. Key contributions include: (i) the first integration of interpretable, rule-based logic into a deep learning pipeline, balancing modeling accuracy with decision transparency; and (ii) substantial reduction in hardware requirements, enabling rapid deployment in high-risk environments. Evaluated on two real-world hydraulic systems, the generated graphs achieve high topological fidelity to ground truth (average F1 score = 0.89), demonstrating strong applicability and engineering viability.
This work addresses key challenges in deploying 3D deep learning on edge devices, including the unstructured nature of point clouds, high computational costs of conventional preprocessing, and performance degradation caused by domain gaps between synthetic CAD models and real-world LiDAR data. To bridge this domain discrepancy, the authors propose a sensor-aware physically simulated LiDAR data generation method. Furthermore, they introduce a deterministic Critical Point Layer (CPL) that enables efficient point cloud compression without requiring distance-based sorting. Integrated with an ARM Cortex-A76-optimized lightweight classification network, the system compresses input point clouds from 1,024 to 40–60 points and achieves real-time inference at approximately 50 FPS on a Raspberry Pi 5, attaining a classification accuracy of 88.36%.