🤖 AI Summary
This study addresses the insufficient acceleration of banded sparse attention matrix multiplication in Longformer by proposing an FPGA-oriented pipelined hardware architecture. The core innovation lies in a novel implicit index storage scheme that eliminates the overhead of conventional sparse formats to enable regularized memory access. This is combined with parallel processing units, a pipelined adder tree, and a dual-path computation strategy to achieve highly efficient operations. Validated on an RFSoC platform operating at 100 MHz, the proposed architecture produces one dot-product result per clock cycle while consuming less than 2.9 W of power and delivering a computational throughput exceeding 100 million operations per second. Consequently, this work presents a highly energy-efficient hardware acceleration solution for structured sparse Transformers.
📝 Abstract
Sparse attention mechanisms have become increasingly important for transformer models processing long input sequences due to their lower computational and memory complexity compared to full self-attention. Longformer achieves this through a sliding-window attention mechanism that produces a structured banded sparse attention matrix. However, existing sparse transformer accelerators primarily target attention generation or unstructured sparsity, leaving sparse matrix--dense matrix multiplication (SpMM) for structured sparse attention largely unexplored. This paper presents a pipelined FPGA architecture for accelerating banded SpMM in Longformer. The proposed design exploits the predictable sparsity pattern of Longformer's attention matrix through a custom row-wise storage scheme with implicit indexing, eliminating the overhead of conventional sparse matrix formats while enabling regular memory accesses. The architecture employs parallel processing elements, pipelined adder trees, and a dual-path computation strategy to maximize throughput and hardware utilization. Implemented in Verilog and evaluated on an RFSoC platform using Vivado 2024.2, the accelerator sustains one complete dot-product result per clock cycle after an initial latency of 11 cycles while maintaining power consumption below 2.9 W. Operating at 100 MHz, the design achieves over 100 million dot-product outputs per second, demonstrating the effectiveness of directly exploiting structured sparsity for sparse transformer acceleration.