🤖 AI Summary
This work addresses the lack of rigorous theoretical analysis regarding the expressive power of Transformers, particularly in approximating general function classes. By establishing an explicit approximation relationship between Transformers and Maxout networks, the study constructs a theoretical framework that reveals key structural properties: the self-attention layer can implement max-type operations, while the feedforward layer performs token-wise affine transformations. For the first time, the paper connects the representational capacity of Transformers to classical approximation theory for feedforward networks, proving that Transformers can approximate Maxout networks with comparable complexity, thereby inheriting the universal approximation capability of ReLU networks. Furthermore, it quantifies how the number of linear regions grows exponentially with depth, elucidating the role of depth in enhancing expressive power.
📝 Abstract
Transformer networks have achieved remarkable empirical success across a wide range of applications, yet their theoretical expressive power remains insufficiently understood. In this paper, we study the expressive capabilities of Transformer architectures. We first establish an explicit approximation of maxout networks by Transformer networks while preserving comparable model complexity. As a consequence, Transformers inherit the universal approximation capability of ReLU networks under similar complexity constraints. Building on this connection, we develop a framework to analyze the approximation of continuous piecewise linear functions by Transformers and quantitatively characterize their expressivity via the number of linear regions, which grows exponentially with depth. Our analysis establishes a theoretical bridge between approximation theory for standard feedforward neural networks and Transformer architectures. It also yields structural insights into Transformers: self-attention layers implement max-type operations, while feedforward layers realize token-wise affine transformations.