What can linear attention learn from nonlinear teachers in-context?

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear capabilities and limitations of linear attention mechanisms in handling nonlinear target functions during in-context learning. To this end, it extends the theoretical framework of linear attention to nonlinear single-index models by integrating asymptotic analysis with Hermite polynomial decomposition, thereby introducing a "nonlinearity-noise equivalence" principle. The findings reveal that such models can only extract the linear Hermite component of the target function. Furthermore, this work quantifies the impact of nonlinear structures on generalization error and rigorously characterizes the phase transition conditions from memorization to generalization. Collectively, these contributions establish a novel theoretical foundation for understanding the underlying learning mechanisms of linear attention.
📝 Abstract
Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers. For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour. We extend this theory to nonlinear single-index targets, $y=f(x^\top w)+\varepsilon $. Our main result establishes a nonlinearity-noise equivalence: linear attention extracts only the linear Hermite component of $f$, while the remaining nonlinear structure contributes to the generalisation error as effective noise. This reduction allows results from the corresponding linear theory to be transferred to nonlinear tasks. We illustrate its implications for finite pretraining data and for the transition from task memorisation to task generalisation as task diversity increases. These results identify a limitation of the reduced linear-attention model and provide a tractable starting point for studying nonlinear in-context learning.
Problem

Research questions and friction points this paper is trying to address.

linear attention
in-context learning
nonlinear single-index targets
generalization error
task generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Linear attention
In-context learning
Nonlinearity-noise equivalence
Single-index model
Hermite component
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.