🤖 AI Summary
This study investigates the boundary between the linear and nonlinear capabilities of the Softmax attention mechanism in Transformers, which remains poorly understood. Employing Gaussian mixture statistical modeling, infinite-prompt asymptotic analysis, and gradient optimization theory, this work examines the asymptotic behavior of attention under supervised classification and denoising tasks. We demonstrate that Softmax attention can precisely learn optimal solutions across diverse statistical tasks, revealing its dual nature: it reduces to a linear mapping for simple tasks, yet transcends linear limitations through query-dependent selection in complex scenarios. Consequently, this research establishes that Softmax attention inherently combines both linear recovery and nonlinear context selection capabilities, providing a rigorous theoretical characterization of its representational power.
📝 Abstract
Softmax attention, at the heart of Transformers, has demonstrated remarkable capabilities. Yet its underlying mechanisms remain only partially understood. Recent theoretical work studies Gaussian prompts, where the infinite-prompt limit reduces softmax attention to a linear map, but also removes the query-dependent selection that distinguishes it from linear attention. This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies. We show that softmax attention can represent and learn, via gradient-based methods, optimal solutions to a range of statistical tasks, including supervised classification and denoising. Our results highlight two complementary capabilities of softmax attention: it can recover linear tasks as effectively as its simpler linear counterpart, while also exploiting query-dependent context selection to solve nonlinear tasks beyond the reach of linear attention.