Gaussian Mixture Attention: Linear-Time Sequence Mixing via Probabilistic Latent Routing

๐Ÿ“… 2026-06-09
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Standard dot-product attention incurs computational and memory bottlenecks in long-context scenarios due to dense pairwise interactions. This work proposes Gaussian Mixture Attention (GMA), which maps queries and keys into a shared latent routing space and implicitly computes similarities through K learnable Gaussian mixture components, while reading from and writing to a K-slot latent memoryโ€”thereby avoiding explicit construction of the Nร—N attention matrix. GMA achieves linear-time sequence mixing via probabilistic latent-variable routing, offering interpretable responsibility assignment, non-negative low-rank approximation, and stable local routing. Its end-to-end differentiable design supports both causal and bidirectional variants, enabling linear memory scaling with fixed K. Empirically, GMA matches standard attention in long-context classification, outperforms several linear methods on WikiText-103 in its causal form, and exhibits broad, semantically aligned component utilization as revealed by responsibility analysis.
๐Ÿ“ Abstract
The dense token-to-token interaction pattern of standard dot-product attention remains a central bottleneck in scaling Transformer architectures to long contexts. We introduce \textbf{Gaussian Mixture Attention (GMA)}, a probabilistic attention-style sequence mixer that replaces explicit pairwise query--key comparison with routing through $K$ learned Gaussian mixture components. Queries and keys are mapped to posterior \textit{responsibility} vectors over a shared latent routing space; their overlap defines an implicit responsibility-space affinity, while values are written into and read from a $K$-slot latent memory. By exploiting the associativity of matrix multiplication, GMA avoids materializing the induced $N\times N$ affinity matrix and instead uses two responsibility matrices whose dominant activation storage scales as $\mathcal{O}(NK)$ rather than $\mathcal{O}(N^2)$ for fixed $K$. We formulate bidirectional and causal variants of GMA, provide an end-to-end differentiable parameterization of the Gaussian mixture components, and analyze its responsibility-modulated gradient structure, constrained non-negative low-rank affinity interpretation, and local routing stability. Empirically, GMA exhibits the intended fixed-$K$ linear memory scaling and is competitive with attention-style baselines on long-context classification, while causal GMA improves over tested linear/random-feature attention variants on WikiText-103 but remains behind optimized causal SDPA and Mamba in the current implementation. Analysis of learned responsibilities further shows broad component usage and moderate alignment with surface-form token categories, supporting GMA as a probabilistic, interpretable, fixed-$K$ linear-time attention-style alternative rather than a universal replacement for optimized softmax attention or state-space models.
Problem

Research questions and friction points this paper is trying to address.

attention mechanism
long-context scaling
computational bottleneck
memory complexity
sequence modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gaussian Mixture Attention
linear-time attention
probabilistic latent routing
responsibility vectors
low-rank affinity
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.