Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling

📅 2025-08-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Transformer’s self-attention incurs O(N²) computational complexity, hindering efficient long-sequence modeling. To address this, we propose Gated Associative Memory (GAM), a sequence modeling architecture with linear time complexity O(N). GAM innovatively integrates local causal convolutions with global parallel associative memory retrieval, and introduces a dual-path gated fusion mechanism that enables dynamic, fully parallel coordination of local and global information—achieved without approximation or sparsification. Unlike prior linear-time models, GAM is implemented from first principles. Experiments on WikiText-2 and TinyStories demonstrate that GAM trains significantly faster than Transformer and Mamba baselines while achieving comparable or superior validation perplexity, confirming its dual advantages in computational efficiency and representational capacity.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: (Computational) Cognitive Architectures

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Large language models for search
📝 Abstract
The Transformer architecture, underpinned by the self-attention mechanism, has become the de facto standard for sequence modeling tasks. However, its core computational primitive scales quadratically with sequence length (O(N^2)), creating a significant bottleneck for processing long contexts. In this paper, we propose the Gated Associative Memory (GAM) network, a novel, fully parallel architecture for sequence modeling that exhibits linear complexity (O(N)) with respect to sequence length. The GAM block replaces the self-attention layer with two parallel pathways: a causal convolution to efficiently capture local, position-dependent context, and a parallel associative memory retrieval mechanism to model global, content-based patterns. These pathways are dynamically fused using a gating mechanism, allowing the model to flexibly combine local and global information for each token. We implement GAM from scratch and conduct a rigorous comparative analysis against a standard Transformer model and a modern linear-time baseline (Mamba) on the WikiText-2 benchmark, as well as against the Transformer on the TinyStories dataset. Our experiments demonstrate that GAM is consistently faster, outperforming both baselines on training speed, and achieves a superior or competitive final validation perplexity across all datasets, establishing it as a promising and efficient alternative for sequence modeling.
Problem

Research questions and friction points this paper is trying to address.

Addresses quadratic complexity bottleneck in Transformer self-attention
Proposes linear-time architecture for efficient long-sequence modeling
Combines local convolution and global memory retrieval pathways
Innovation

Methods, ideas, or system contributions that make the work stand out.

Linear complexity O(N) parallel architecture
Combines causal convolution and associative memory
Dynamic gating fuses local and global information
🔎 Similar Papers
No similar papers found.
Independent Researcher
R
Rishiraj Acharya
Independent Researcher