What can linear attention learn from nonlinear teachers in-context?
This study addresses the unclear capabilities and limitations of linear attention mechanisms in handling nonlinear target functions during in-context learning. To this end, it extends the theoretical framework of linear attention to nonlinear single-index models by integrating asymptotic analysis with Hermite polynomial decomposition, thereby introducing a "nonlinearity-noise equivalence" principle. The findings reveal that such models can only extract the linear Hermite component of the target function. Furthermore, this work quantifies the impact of nonlinear structures on generalization error and rigorously characterizes the phase transition conditions from memorization to generalization. Collectively, these contributions establish a novel theoretical foundation for understanding the underlying learning mechanisms of linear attention.