Efficient Event-based Semantic Segmentation with Spike-driven Lightweight Transformer-based Networks

πŸ“… 2024-12-17
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
To address the high computational cost, frame-based dependency, and excessive power consumption of conventional artificial neural networks (ANNs) in event-camera semantic segmentation, this paper proposes SLTNetβ€”a lightweight spike-driven Transformer network. Its core innovations are the spike-driven convolutional block (SCB) and the binary-masked spike Transformer block (STB), jointly enabling a single-branch, highly efficient architecture that preserves the ultra-low power advantage of spiking neural networks (SNNs) while significantly enhancing long-range contextual modeling. On the DDD17 and DSEC-Semantic benchmarks, SLTNet achieves mIoU improvements of 9.06% and 9.39% over state-of-the-art SNN methods, respectively, reduces power consumption by 4.58Γ—, and attains real-time inference at 114 FPS. This work establishes a new paradigm for edge-deployable, real-time, low-power event-based semantic segmentation.

Technology Category

Cognitive Modeling & Cognitive Systems: Neural Spike CodingMachine Learning: Learning on the Edge & Model CompressionComputer Vision: Segmentation

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSecurity and Privacy: Large-scale security measurements
πŸ“ Abstract
Event-based semantic segmentation has great potential in autonomous driving and robotics due to the advantages of event cameras, such as high dynamic range, low latency, and low power cost. Unfortunately, current artificial neural network (ANN)-based segmentation methods suffer from high computational demands, the requirements for image frames, and massive energy consumption, limiting their efficiency and application on resource-constrained edge/mobile platforms. To address these problems, we introduce SLTNet, a spike-driven lightweight transformer-based network designed for event-based semantic segmentation. Specifically, SLTNet is built on efficient spike-driven convolution blocks (SCBs) to extract rich semantic features while reducing the model's parameters. Then, to enhance the long-range contextural feature interaction, we propose novel spike-driven transformer blocks (STBs) with binary mask operations. Based on these basic blocks, SLTNet employs a high-efficiency single-branch architecture while maintaining the low energy consumption of the Spiking Neural Network (SNN). Finally, extensive experiments on DDD17 and DSEC-Semantic datasets demonstrate that SLTNet outperforms state-of-the-art (SOTA) SNN-based methods by at most 9.06% and 9.39% mIoU, respectively, with extremely 4.58x lower energy consumption and 114 FPS inference speed. Our code is open-sourced and available at https://github.com/longxianlei/SLTNet-v1.0.
Problem

Research questions and friction points this paper is trying to address.

High computational demands in event-based semantic segmentation.
Energy inefficiency in current ANN-based segmentation methods.
Limited application on resource-constrained edge/mobile platforms.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spike-driven lightweight transformer-based network
Efficient spike-driven convolution blocks (SCBs)
Novel spike-driven transformer blocks (STBs)
πŸ’Ό Related Jobs
No related jobs found.
Chongqing University | Institute of Automation, Chinese Academy of Sciences
X
Xiaxin Zhu
College of Computer Science, Chongqing University, Chongqing, China
F
Fangming Guo
College of Computer Science, Chongqing University, Chongqing, China
X
Xianlei Long
College of Computer Science, Chongqing University, Chongqing, China
Qingyi Gu
Qingyi Gu
Institute of Automation, Chinese Academy of Sciences
High-speed visioncell analysis
C
Chao Chen
College of Computer Science, Chongqing University, Chongqing, China
Fuqiang Gu
Fuqiang Gu
Professor, College of Computer Science, Chongqing University
Indoor positioningdeep learningrobotsSLAMbrain-like computing