🤖 AI Summary
Transformer’s self-attention incurs O(N²) computational complexity, hindering efficient long-sequence modeling. To address this, we propose Gated Associative Memory (GAM), a sequence modeling architecture with linear time complexity O(N). GAM innovatively integrates local causal convolutions with global parallel associative memory retrieval, and introduces a dual-path gated fusion mechanism that enables dynamic, fully parallel coordination of local and global information—achieved without approximation or sparsification. Unlike prior linear-time models, GAM is implemented from first principles. Experiments on WikiText-2 and TinyStories demonstrate that GAM trains significantly faster than Transformer and Mamba baselines while achieving comparable or superior validation perplexity, confirming its dual advantages in computational efficiency and representational capacity.
📝 Abstract
The Transformer architecture, underpinned by the self-attention mechanism, has become the de facto standard for sequence modeling tasks. However, its core computational primitive scales quadratically with sequence length (O(N^2)), creating a significant bottleneck for processing long contexts. In this paper, we propose the Gated Associative Memory (GAM) network, a novel, fully parallel architecture for sequence modeling that exhibits linear complexity (O(N)) with respect to sequence length. The GAM block replaces the self-attention layer with two parallel pathways: a causal convolution to efficiently capture local, position-dependent context, and a parallel associative memory retrieval mechanism to model global, content-based patterns. These pathways are dynamically fused using a gating mechanism, allowing the model to flexibly combine local and global information for each token. We implement GAM from scratch and conduct a rigorous comparative analysis against a standard Transformer model and a modern linear-time baseline (Mamba) on the WikiText-2 benchmark, as well as against the Transformer on the TinyStories dataset. Our experiments demonstrate that GAM is consistently faster, outperforming both baselines on training speed, and achieves a superior or competitive final validation perplexity across all datasets, establishing it as a promising and efficient alternative for sequence modeling.