On the Expressive Power of Transformers for Maxout Networks and Continuous Piecewise Linear Functions

📅 2026-03-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of rigorous theoretical analysis regarding the expressive power of Transformers, particularly in approximating general function classes. By establishing an explicit approximation relationship between Transformers and Maxout networks, the study constructs a theoretical framework that reveals key structural properties: the self-attention layer can implement max-type operations, while the feedforward layer performs token-wise affine transformations. For the first time, the paper connects the representational capacity of Transformers to classical approximation theory for feedforward networks, proving that Transformers can approximate Maxout networks with comparable complexity, thereby inheriting the universal approximation capability of ReLU networks. Furthermore, it quantifies how the number of linear regions grows exponentially with depth, elucidating the role of depth in enhancing expressive power.

Technology Category

Machine Learning: Deep Neural Architectures and Foundation ModelsNatural Language Processing: Learning & Optimization for NLPComputer Vision: Representation Learning for Vision

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Web query analysis, representation and understanding
📝 Abstract
Transformer networks have achieved remarkable empirical success across a wide range of applications, yet their theoretical expressive power remains insufficiently understood. In this paper, we study the expressive capabilities of Transformer architectures. We first establish an explicit approximation of maxout networks by Transformer networks while preserving comparable model complexity. As a consequence, Transformers inherit the universal approximation capability of ReLU networks under similar complexity constraints. Building on this connection, we develop a framework to analyze the approximation of continuous piecewise linear functions by Transformers and quantitatively characterize their expressivity via the number of linear regions, which grows exponentially with depth. Our analysis establishes a theoretical bridge between approximation theory for standard feedforward neural networks and Transformer architectures. It also yields structural insights into Transformers: self-attention layers implement max-type operations, while feedforward layers realize token-wise affine transformations.
Problem

Research questions and friction points this paper is trying to address.

expressive power
Transformers
maxout networks
continuous piecewise linear functions
approximation theory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformers
maxout networks
continuous piecewise linear functions
expressive power
universal approximation
🔎 Similar Papers
No similar papers found.
L
Linyan Gu
School of Mathematics, Sun Yat-sen University
L
Lihua Yang
School of Mathematics, Sun Yat-sen University
F
Feng Zhou
School of Information Science, Guangdong University of Finance and Economics