An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer

📅 2025-01-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Spiking Transformer architectures suffer from high power consumption due to dense, continuous computation. Method: This work proposes a hardware accelerator tailored for sparse spike signals, featuring a novel spiking self-attention unit that supports dual-spike inputs, integrated with spike-based positional encoding and a sparse activation routing mechanism to enable end-to-end event-driven computation. Crucially, it skips non-spiking positions entirely—performing linear transformations, pooling, and self-attention only on active spikes. Contribution/Results: Experimental evaluation demonstrates a 13.24× throughput improvement and 1.33× energy efficiency gain over state-of-the-art SNN accelerators. The design significantly reduces redundant computation and latency, establishing an efficient, low-power hardware paradigm for large-scale spiking neural networks.

Technology Category

Cognitive Modeling & Cognitive Systems: Neural Spike CodingMachine Learning: Hardware-aware MLSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSecurity and Privacy: Large-scale security measurementsGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphs
📝 Abstract
Recently, large models, such as Vision Transformer and BERT, have garnered significant attention due to their exceptional performance. However, their extensive computational requirements lead to considerable power and hardware resource consumption. Brain-inspired computing, characterized by its spike-driven methods, has emerged as a promising approach for low-power hardware implementation. In this paper, we propose an efficient sparse hardware accelerator for Spike-driven Transformer. We first design a novel encoding method that encodes the position information of valid activations and skips non-spike values. This method enables us to use encoded spikes for executing the calculations of linear, maxpooling and spike-driven self-attention. Compared with the single spike input design of conventional SNN accelerators that primarily focus on convolution-based spiking computations, the specialized module for spike-driven self-attention is unique in its ability to handle dual spike inputs. By exclusively utilizing activated spikes, our design fully exploits the sparsity of Spike-driven Transformer, which diminishes redundant operations, lowers power consumption, and minimizes computational latency. Experimental results indicate that compared to existing SNNs accelerators, our design achieves up to 13.24$ imes$ and 1.33$ imes$ improvements in terms of throughput and energy efficiency, respectively.
Problem

Research questions and friction points this paper is trying to address.

Hardware Accelerator
Large-scale Model Processing
Energy Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synapse-driven Transformer
Sparse Encoding
Hardware Accelerator
Z
Zhengke Li
School of Integrated Circuits, Sun Yat-Sen University, China; School of Electronic Science and Engineering, Nanjing University, China
W
W. Mao
School of Integrated Circuits, Sun Yat-Sen University, China; School of Electronic Science and Engineering, Nanjing University, China
Siyu Zhang
Siyu Zhang
4DV.ai
Computer Vision
Q
Qiwei Dong
School of Electronic Science and Engineering, Nanjing University, China
Zhongfeng Wang
Zhongfeng Wang
Nanjing University
VLSIFECDSPMIMONeural Network