Switching Linear Attention

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fundamental trade-off between the limited expressiveness of linear attention and the prohibitive memory overhead of Softmax attention by proposing the SwiLA layer. Grounded in a test-time regression framework, this method formulates state updates as an online Expectation-Maximization (EM) algorithm for mixture-of-linear-regressions models. It introduces a novel mechanism that integrates fixed-size recurrent states with dynamic output dimension selection, enabling adaptive component selection to effectively balance efficiency and performance. Experimental evaluations demonstrate that SwiLA achieves strong performance on associative recall and language modeling benchmarks, surpassing standard Softmax attention in certain scenarios. Ultimately, this work successfully reconciles computational efficiency with model expressiveness, offering a compelling alternative to conventional attention mechanisms.
📝 Abstract
Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning. Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequence length, limiting its scalability. Linear attention enables efficient recurrent computation with a constant memory footprint, yet its reduced expressivity often yields inferior modeling performance. We introduce Switching Linear Attention (SwiLA), a novel sequence layer that bridges this gap by enhancing representational capacity while retaining the fixed-size recurrent state of linear attention. We derive the SwiLA recurrence from the test-time regression framework, casting the state update rule as online expectation-maximization in a mixture of linear regressions model. At test time, each output dimension dynamically selects among multiple linear attention components based on the input. Across associative recall, in-context language learning, and language modeling benchmarks, SwiLA shows strong performance and narrows the gap to softmax attention, even surpassing it in several settings.
Problem

Research questions and friction points this paper is trying to address.

Linear Attention
Sequence Modeling
Efficient Inference
Expressivity
Softmax Attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Switching Linear Attention
Test-time Regression
Online Expectation-Maximization
Mixture of Linear Regressions
Efficient Inference
🔎 Similar Papers
No similar papers found.