A Pipelined FPGA Architecture for Banded Sparse Matrix Dense Matrix Multiplication in Longformer

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient acceleration of banded sparse attention matrix multiplication in Longformer by proposing an FPGA-oriented pipelined hardware architecture. The core innovation lies in a novel implicit index storage scheme that eliminates the overhead of conventional sparse formats to enable regularized memory access. This is combined with parallel processing units, a pipelined adder tree, and a dual-path computation strategy to achieve highly efficient operations. Validated on an RFSoC platform operating at 100 MHz, the proposed architecture produces one dot-product result per clock cycle while consuming less than 2.9 W of power and delivering a computational throughput exceeding 100 million operations per second. Consequently, this work presents a highly energy-efficient hardware acceleration solution for structured sparse Transformers.
📝 Abstract
Sparse attention mechanisms have become increasingly important for transformer models processing long input sequences due to their lower computational and memory complexity compared to full self-attention. Longformer achieves this through a sliding-window attention mechanism that produces a structured banded sparse attention matrix. However, existing sparse transformer accelerators primarily target attention generation or unstructured sparsity, leaving sparse matrix--dense matrix multiplication (SpMM) for structured sparse attention largely unexplored. This paper presents a pipelined FPGA architecture for accelerating banded SpMM in Longformer. The proposed design exploits the predictable sparsity pattern of Longformer's attention matrix through a custom row-wise storage scheme with implicit indexing, eliminating the overhead of conventional sparse matrix formats while enabling regular memory accesses. The architecture employs parallel processing elements, pipelined adder trees, and a dual-path computation strategy to maximize throughput and hardware utilization. Implemented in Verilog and evaluated on an RFSoC platform using Vivado 2024.2, the accelerator sustains one complete dot-product result per clock cycle after an initial latency of 11 cycles while maintaining power consumption below 2.9 W. Operating at 100 MHz, the design achieves over 100 million dot-product outputs per second, demonstrating the effectiveness of directly exploiting structured sparsity for sparse transformer acceleration.
Problem

Research questions and friction points this paper is trying to address.

Sparse Attention
Banded SpMM
Longformer
FPGA Accelerator
Structured Sparsity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pipelined FPGA Architecture
Banded SpMM
Longformer
Implicit Indexing
Structured Sparsity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Phillip Pramberger
Department of Electronic and Electrical Engineering, Trinity College Dublin, Ireland
A
Athanasios Tziouvaras
Department of Electrical and Computer Engineering, University of Thessaly, Greece
Shreejith Shanker
Shreejith Shanker
Assistant Professor, Trinity College Dublin
Reconfigurable ComputingFPGAsEmbedded SystemsComputer ArchitectureMachine Learning
George Floros
George Floros
Assistant Professor, Electronic & Electrical Engineering, Trinity College Dublin
VLSIEDAReliabilityCircuit Simulation