Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension

📅 2026-04-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing spiking Transformers lack theoretical grounding, leading to a disconnect between their representational capacity and empirical performance. This work establishes the first universal approximation theory for spiking self-attention, proving its ability to approximate any continuous permutation-equivariant function, and introduces a design criterion based on effective dimensionality. By integrating Leaky Integrate-and-Fire neurons with rate–distortion theory and information-theoretic analysis, the method incorporates lateral inhibition to implement softmax normalization and derives a tight lower bound on the required number of spikes, determined by the input-dependent effective dimension. Experiments across multiple spiking Transformer architectures and vision–language benchmarks validate the theoretical predictions (R² = 0.97, p < 0.001), demonstrating that high performance can be achieved with as few as four timesteps.

Technology Category

Cognitive Modeling & Cognitive Systems: Neural Spike CodingMachine Learning: Information TheoryNatural Language Processing: Learning & Optimization for NLP

Application Category

Search and Retrieval-Augmented AI: Web query analysis, representation and understandingSecurity and Privacy: Large-scale security measurementsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Spiking transformers achieve competitive accuracy with conventional transformers while offering $38$-$57\times$ energy efficiency on neuromorphic hardware, yet no theoretical framework guides their design. This paper establishes the first comprehensive expressivity theory for spiking self-attention. We prove that spiking attention with Leaky Integrate-and-Fire neurons is a universal approximator of continuous permutation-equivariant functions, providing explicit spike circuit constructions including a novel lateral inhibition network for softmax normalization with proven $O(1/\sqrt{T})$ convergence. We derive tight spike-count lower bounds via rate-distortion theory: $\varepsilon$-approximation requires $Ω(L_f^2 nd/\varepsilon^2)$ spikes, with rigorous information-theoretic derivation. Our key insight is input-dependent bounds using measured effective dimensions ($d_{\text{eff}}=47$--$89$ for CIFAR/ImageNet), explaining why $T=4$ timesteps suffice despite worst-case $T \geq 10{,}000$ predictions. We provide concrete design rules with calibrated constants ($C=2.3$, 95\% CI: $[1.9, 2.7]$). Experiments on Spikformer, QKFormer, and SpikingResformer across vision and language benchmarks validate predictions with $R^2=0.97$ ($p<0.001$). Our framework provides the first principled foundation for neuromorphic transformer design.
Problem

Research questions and friction points this paper is trying to address.

spiking transformers
theory-practice gap
expressivity theory
neuromorphic hardware
effective dimension
Innovation

Methods, ideas, or system contributions that make the work stand out.

spiking transformers
effective dimension
universal approximation
lateral inhibition
rate-distortion theory
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Dongxin Guo
The University of Hong Kong, Hong Kong, China
J
Jikun Wu
Brain Investing Limited, Hong Kong, China
S
Siu Ming Yiu
The University of Hong Kong, Hong Kong, China