🤖 AI Summary
This study addresses the limitation of gated linear attention, which is constrained by a single temporal resolution and thus struggles to simultaneously capture local syntax and long-range semantics. To overcome this, we propose a multi-scale gated linear attention mechanism that assigns distinct attention heads to multiple temporal resolutions and introduces an input-dependent learnable fusion layer to dynamically recombine their outputs. Grounded in the principles of multi-scale state space models, this design substantially expands the effective memory capacity without increasing the state dimension of individual heads. Experimental evaluations demonstrate that the proposed approach achieves superior performance on language modeling and recall tasks, reducing perplexity by 9.5% and improving recall performance by 18.9%.
📝 Abstract
Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously encode local syntactic patterns and long-range semantic structure, creating a representational bottleneck that gating alone is insufficient to resolve. We introduce Multi-Scale Gated Linear Attention (MS-GLA), which addresses this by distributing attention heads across multiple temporal resolutions. Coarser resolutions pool longer token spans naturally specializing toward long-range dependencies, while finer head groups retain sensitivity to local syntactic structure. A learnable, input-dependent fusion layer dynamically recombines head group outputs at each timestep, expanding effective memory capacity without increasing per-head state size. This multi-resolution decomposition draws on principles from Multi-Scale State-Space Models (MS-SSM), adapting them to the gated linear attention setting. We evaluate MS-GLA on language modeling, recall-intensive tasks, and long-context generalization. Across all settings, MS-GLA consistently achieves higher accuracy and lower perplexity than GLA at matched parameter counts, with up to 18.9% improvement on recall-intensive tasks and 9.5% lower average perplexity on language modeling benchmarks, validating multi-temporal resolution decomposition as a principled and effective extension of Gated Linear Attention.