Generalization error bounds for two-layer neural networks with Lipschitz loss function

📅 2026-04-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates generalization error bounds for two-layer neural networks without assuming boundedness of the loss function. By integrating Wasserstein distance to quantify the discrepancy between the true data distribution and the empirical measure, empirical process theory, and moment estimates of stochastic gradient methods, the authors derive the first explicit, computable, and dimension-dependent generalization bounds that apply to both independent and non-independent test data settings. In the independent case, they achieve a dimension-free rate of $O(n^{-1/2})$, while for non-independent data, they obtain a bound of order $O(n^{-1/(d_{\text{in}} + d_{\text{out}})})$. Numerical experiments corroborate the theoretical findings.

Technology Category

Machine Learning: Learning with ManifoldsNatural Language Processing: Learning & Optimization for NLPSearch and Optimization: Non-convex Optimization

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
We derive generalization error bounds for the training of two-layer neural networks without assuming boundedness of the loss function, using Wasserstein distance estimates on the discrepancy between a probability distribution and its associated empirical measure, together with moment bounds for the associated stochastic gradient method. In the case of independent test data, we obtain a dimension-free rate of order $O(n^{-1/2} )$ on the $n$-sample generalization error, whereas without independence assumption, we derive a bound of order $O(n^{-1 / ( d_{\rm in}+d_{\rm out} )} )$, where $d_{\rm in}$, $d_{\rm out}$ denote input and output dimensions. Our bounds and their coefficients can be explicitly computed prior to the training of the model, and are confirmed by numerical simulations.
Problem

Research questions and friction points this paper is trying to address.

generalization error
two-layer neural networks
Lipschitz loss function
Wasserstein distance
empirical measure
Innovation

Methods, ideas, or system contributions that make the work stand out.

generalization error bounds
two-layer neural networks
Wasserstein distance
Lipschitz loss
dimension-free rate
💼 Related Jobs
No related jobs found.
J
Jiang Yu Nguwi
Division of Mathematical Sciences, School of Physical and Mathematical Sciences, Nanyang Technological University
Nicolas Privault
Nicolas Privault
Nanyang Technological University
Stochastic Analysis