Parallel Layer Normalization for Universal Approximation

📅 2025-05-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the long-standing theoretical gap in the Universal Approximation Theorem (UAT)—its neglect of normalization layers—by incorporating Layer Normalization into the UAT framework for the first time. Specifically, it introduces Parallel Layer Normalization (PLN), a novel architectural primitive that jointly performs normalization and nonlinear activation. Leveraging functional approximation theory and neural representational capacity modeling, we rigorously prove that infinitely wide networks composed solely of PLN and linear layers are universal approximators; moreover, we precisely characterize the minimum neuron count required for a single-hidden-layer PLN network to approximate $L$-Lipschitz functions. Theoretically, PLN achieves approximation efficiency comparable to—or even exceeding—that of conventional activation functions. Empirically, substituting standard LayerNorm with PLN in Transformers yields significant performance gains. This work establishes the first UAT-based theoretical foundation for normalized neural networks and reveals the intrinsic representational power of normalization operations.

Technology Category

Machine Learning: Deep Neural Architectures and Foundation ModelsNatural Language Processing: (Large) Language ModelsComputer Vision: Large Vision Models

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Universal approximation theorem (UAT) is a fundamental theory for deep neural networks (DNNs), demonstrating their powerful representation capacity to represent and approximate any function. The analyses and proofs of UAT are based on traditional network with only linear and nonlinear activation functions, but omitting normalization layers, which are commonly employed to enhance the training of modern networks. This paper conducts research on UAT of DNNs with normalization layers for the first time. We theoretically prove that an infinitely wide network -- composed solely of parallel layer normalization (PLN) and linear layers -- has universal approximation capacity. Additionally, we investigate the minimum number of neurons required to approximate $L$-Lipchitz continuous functions, with a single hidden-layer network. We compare the approximation capacity of PLN with traditional activation functions in theory. Different from the traditional activation functions, we identify that PLN can act as both activation function and normalization in deep neural networks at the same time. We also find that PLN can improve the performance when replacing LN in transformer architectures, which reveals the potential of PLN used in neural architectures.
Problem

Research questions and friction points this paper is trying to address.

Proving universal approximation capacity of networks with parallel layer normalization
Determining minimum neurons needed for approximating L-Lipchitz functions
Comparing PLN's approximation capacity with traditional activation functions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parallel Layer Normalization enables universal approximation
PLN combines activation and normalization functions
PLN improves performance in transformer architectures
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yunhao Ni
SKLSDE, School of Artificial Intelligence, Beihang University, Beijing, China
Y
Yuhe Liu
SKLSDE, School of Artificial Intelligence, Beihang University, Beijing, China
Wenxin Sun
Wenxin Sun
University of Liverpool, Xi’an Jiaotong-liverpool University
human-computer interaction
Y
Yitong Tang
SKLSDE, School of Artificial Intelligence, Beihang University, Beijing, China
Y
Yuxin Guo
SKLSDE, School of Artificial Intelligence, Beihang University, Beijing, China
P
Peilin Feng
SKLSDE, School of Artificial Intelligence, Beihang University, Beijing, China
W
Wenjun Wu
SKLSDE, School of Artificial Intelligence, Beihang University, Beijing, China; Hangzhou International Innovation Institute, Beihang University, Hangzhou, China
L
Lei Huang
SKLSDE, School of Artificial Intelligence, Beihang University, Beijing, China; Hangzhou International Innovation Institute, Beihang University, Hangzhou, China