HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference

📅 2024-11-02
🏛️ Micro
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address two critical bottlenecks in embedded real-time point cloud processing—memory-intensive preprocessing (e.g., downsampling) and data-structure-dependent inference—this paper proposes an end-to-end heterogeneous acceleration architecture. We introduce Octree-Indexed Sampling, a spatially aware preprocessing method enabling efficient, memory-frugal downsampling; and Voxel-Expanded Gathering, a novel data structuring mechanism coupled with a custom-designed Deep Learning Accelerator (DLA) expansion unit. The entire pipeline is co-optimized on an Intel Programmable Acceleration Card (PAC) platform integrating Xeon CPUs and FPGAs. Experimental evaluation demonstrates that our approach achieves 1.3–10.2×, 2.2–16.5×, and 6.4–21× inference speedup over PointACC, Mesorasi, and Jetson NX, respectively. Moreover, the end-to-end latency meets the real-time requirements of raw 3D sensor data streams, enabling practical deployment in resource-constrained embedded systems.

Technology Category

Machine Learning: Hardware-aware MLComputer Vision: 3D Computer VisionSearch and Optimization: Learning to Search

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Data management and stream processing for Web, mobile and wireless applicationsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Point cloud is an important type of geometric data structure for many embedded applications such as autonomous driving and augmented reality. Current Point Cloud Networks (PCNs) have proven to achieve great success in using inference to perform point cloud analysis, including object part segmentation, shape classification, and so on. However, point cloud applications on the computing edge require more than just the inference step. They require an end-to-end (E2E) processing of the point cloud workloads: pre-processing of raw data, input preparation, and inference to perform point cloud analysis. Current PCN approaches to support end-to-end processing of point cloud workload cannot meet the real-time latency requirement on the edge, i.e., the ability of the AI service to keep up with the speed of raw data generation by 3D sensors. Latency for end-to-end processing of the point cloud workloads stems from two reasons: memory-intensive down-sampling in the pre-processing phase and the data structuring step for input preparation in the inference phase. In this paper, we present HgPCN, an end-to-end heterogeneous architecture for real-time embedded point cloud applications. In HgPCN, we introduce two novel methodologies based on spatial indexing to address the two identified bottlenecks. In the Pre-processing Engine of HgPCN, an Octree-Indexed-Sampling method is used to optimize the memory-intensive down-sampling bottleneck of the pre-processing phase. In the Inference Engine, HgPCN extends a commercial DLA with a customized Data Structuring Unit which is based on a Voxel-Expanded Gathering method to fundamentally reduce the workload of the data structuring step in the inference phase. The initial prototype of HgPCN has been implemented on an Intel PAC (Xeon+FPGA) platform. Four commonly available point cloud datasets were used for comparison, running on three baseline devices: Intel Xeon W-2255, Nvidia Xavier NX Jetson GPU, and Nvidia 4060ti GPU. These point cloud datasets were also run on two existing PCN accelerators for comparison: PointACC and Mesorasi. Our results show that for the inference phase, depending on the dataset size, HgPCN achieves speedup from 1.3× To 10.2× vs. PointACC, 2.2× To 16.5× vs. Mesorasi, and 6.4× To 21× vs. Jetson NX GPU. Along with optimization of the memory-intensive down-sampling bottleneck in pre-processing phase, the overall latency shows that HgPCN can reach the real-time requirement by providing end-to-end service with keeping up with the raw data generation rate.
Problem

Research questions and friction points this paper is trying to address.

Point Cloud Networks
Real-time Processing
Autonomous Driving and Augmented Reality
Innovation

Methods, ideas, or system contributions that make the work stand out.

HgPCN Architecture
Octree Index Sampling
Voxel Expansion Aggregation
Y
Yiming Gao
University of Florida
C
Chao Jiang
University of Florida
W
Wesley Piard
University of Florida
X
Xiangru Chen
University of Florida
B
Bhavesh Patel
Dell EMC
Herman Lam
Herman Lam
University of Florida