Score
Designs, implements, and trains transformer-style neural networks whose units communicate with discrete spikes, including spike-based attention modules and spike-driven operations. This includes building encoders and decoders to represent inputs as spike trains, choosing or developing spike-aware learning rules and optimizers, and analyzing sparse event-driven activations, inference activity, and computational efficiency.
Current spiking Transformers lack a principled understanding of the intrinsic signal characteristics of event data, focusing predominantly on architectural modifications while neglecting underlying mechanistic principles. This work is the first to reveal—through a frequency-domain analysis—that spiking Transformers inherently behave as high-pass filters, leading to insufficient extraction of critical mid-frequency features from sparse and noisy event streams. To address this, we propose SpikePool: a max-pooling–based attention mechanism that implements selective band-pass filtering—preserving informative high-frequency dynamics while suppressing high-frequency noise. SpikePool bridges theoretical interpretability and engineering efficacy. It achieves state-of-the-art performance on event-camera classification and detection benchmarks, while reducing training and inference time by 42.5% and 32.8%, respectively. The method significantly enhances the robustness and computational efficiency of spiking Transformers.
Spiking Transformer architectures suffer from high power consumption due to dense, continuous computation. Method: This work proposes a hardware accelerator tailored for sparse spike signals, featuring a novel spiking self-attention unit that supports dual-spike inputs, integrated with spike-based positional encoding and a sparse activation routing mechanism to enable end-to-end event-driven computation. Crucially, it skips non-spiking positions entirely—performing linear transformations, pooling, and self-attention only on active spikes. Contribution/Results: Experimental evaluation demonstrates a 13.24× throughput improvement and 1.33× energy efficiency gain over state-of-the-art SNN accelerators. The design significantly reduces redundant computation and latency, establishing an efficient, low-power hardware paradigm for large-scale spiking neural networks.
Existing ANN-to-SNN conversion methods struggle to handle the complex nonlinear operations in Transformers and rely heavily on post-training fine-tuning, hindering efficient deployment of spiking Transformers. This paper proposes a training-free ANN-to-SNN conversion framework. We introduce Multi-Basis Exponential (MBE) spiking neurons that accurately approximate key nonlinearities—including Softmax, LayerNorm, and GeLU—without weight updates or architectural modifications, enabling plug-and-play integration. Coupled with an event-driven execution paradigm and a Multi-basis encoding strategy, our framework supports mainstream architectures such as ViT, RoBERTa, and GPT-2. Evaluated across computer vision (CV), natural language understanding (NLU), and natural language generation (NLG) tasks, it achieves near-lossless accuracy (average degradation <0.5%), reduces inference latency by 3.2–5.8×, and significantly improves energy efficiency. Our approach establishes a novel paradigm for scalable, high-performance spiking Transformer deployment.
Spiking Transformer models—hybrid architectures combining spiking neural networks (SNNs) and Transformers—exhibit low energy efficiency on general-purpose hardware, while existing neuromorphic or processing-in-memory (PIM) accelerators struggle to handle their high spatio-temporal sparsity and complex operations. Method: This paper proposes a hardware acceleration engine for event-driven visual inference, featuring an analog-digital hybrid PIM architecture, spatio-temporally sparse-aware dataflow, dynamic optimization mechanisms (including layer skipping and timestep reduction), and a Bayesian optimization–driven algorithm–microarchitecture co-design methodology. Contribution/Results: Evaluated on ImageNet, CIFAR-10 DVS, and DVSGesture benchmarks, the accelerator achieves up to 467× and 1.86× higher energy efficiency than edge GPUs and state-of-the-art PIM accelerators, respectively, while maintaining state-of-the-art accuracy.
To address the challenge of simultaneously achieving biological plausibility, hardware efficiency, and competitive performance in spiking neural networks (SNNs) for general supervised classification, this paper proposes a columnar hierarchical SNN architecture tailored for classification. It employs intra-class-difference-driven columnar organization—each column represents a discriminative subcategory—and adopts an all-spiking signal flow with functionally specialized neurons. A biologically grounded learning mechanism is introduced, integrating local anti-Hebbian plasticity with dopamine neuromodulation to replace backpropagation entirely. The method unifies model-driven reinforcement learning with a state-proximity evaluation framework, enabling end-to-end training directly in the spike domain. Experiments demonstrate that the architecture achieves high accuracy and strong generalization across multiple benchmark classification tasks, while exhibiting exceptional compatibility with low-power neuromorphic hardware. This work establishes a novel paradigm for practical, deployable SNNs.
This work addresses the high energy consumption of conventional Transformers in natural language processing, which limits their applicability in energy-constrained scenarios. The study proposes the first fully spiking neural network (SNN)-based Transformer decoder tailored for NLP tasks, introducing a trainable SNN decoder architecture, SNN-compatible normalization techniques, and multiple text-to-spike embedding strategies. The authors systematically evaluate the impact of module substitution, residual connections, and normalization on model performance. By extending SNN Transformers beyond vision-only applications or encoder-only designs, the proposed model maintains functional equivalence to its artificial neural network (ANN) counterpart while achieving a theoretical energy reduction of 87%–93% compared to standard ANN baselines.
Existing theoretical analyses of spiking neural networks (SNNs) lack quantitative characterizations of their universal representational capacity as sequence-to-sequence processors over spike trains. Method: We propose a function approximation framework grounded in spike-train modeling, integrating constructive weight design with rigorous temporal complexity quantification. Contribution/Results: We establish the first constructively provable, near-optimal universal approximation theorem for naturally spikable function classes—i.e., functions admitting efficient spike-based realization. Theorematically, SNNs achieve near-optimal complexity in both neuron count and synaptic weight count. They exhibit significant representational advantages for sparse inputs, low-order temporal functions, and composite functions. Moreover, our analysis provides rigorous theoretical foundations for modular deep SNN architectures and downstream tasks such as spike-sequence classification.
Existing spiking Transformers lack theoretical grounding, leading to a disconnect between their representational capacity and empirical performance. This work establishes the first universal approximation theory for spiking self-attention, proving its ability to approximate any continuous permutation-equivariant function, and introduces a design criterion based on effective dimensionality. By integrating Leaky Integrate-and-Fire neurons with rate–distortion theory and information-theoretic analysis, the method incorporates lateral inhibition to implement softmax normalization and derives a tight lower bound on the required number of spikes, determined by the input-dependent effective dimension. Experiments across multiple spiking Transformer architectures and vision–language benchmarks validate the theoretical predictions (R² = 0.97, p < 0.001), demonstrating that high performance can be achieved with as few as four timesteps.
This work addresses the inefficiencies of conventional neural networks in data and energy consumption, as well as the limited availability of effective learning algorithms for spiking neural networks (SNNs). To bridge this gap, the authors propose Spark, a modular SNN framework that constructs end-to-end models by composing simple plasticity-based components, enabling continuous, batch-free learning. Spark integrates seamlessly with traditional machine learning pipelines while incorporating biologically inspired continual learning mechanisms, substantially improving data efficiency and practical applicability. The framework’s effectiveness is demonstrated on the sparse-reward CartPole task, where it successfully learns in a continuous setting, highlighting the promise of modular design for efficient SNN training.
Existing ANN-to-SNN conversion methods struggle to efficiently implement nonlinear operators in Transformers—such as Softmax, SiLU, and normalization layers—due to their reliance on division, exponentiation, and ℓ²-norm operations, which are incompatible with leaky integrate-and-fire (LIF) neuron dynamics. This work proposes a plug-and-play spiking operator framework that, for the first time, systematically decomposes these nonlinearities into three fundamental primitives. By leveraging population coding with LIF neurons and lightweight shift-scale operations, the framework achieves floating-point-free, training-free, and spike-friendly approximations. Its modular design seamlessly integrates into existing ANN-to-SNN pipelines without requiring fine-tuning, supporting mainstream nonlinear functions while preserving model performance across various Transformer architectures with less than 1% accuracy degradation upon operator replacement.