🤖 AI Summary
This work investigates the conditions and structural characteristics under which shallow neural networks achieve mathematical optimality in Banach function spaces.
Method: For networks employing multivariate nonlinear activations—including ReLU, norm-based, and radial basis functions—we construct a novel Banach space grounded in k-plane transformations and sparse norms, and establish its reproducing kernel structure.
Contribution/Results: We prove that the optimal approximation in this space is necessarily realized by a multivariate nonlinear neural architecture featuring skip connections, orthogonally normalized weights, and multi-index parameterization. This work establishes, for the first time, a rigorous correspondence between multivariate nonlinear neural networks and Banach-space optimality, unifying diverse classical activation functions within a single variational framework. It yields a universal representation theorem and delivers an analytic characterization of the optimal network structure. Moreover, it provides a rigorous theoretical foundation for regularity analysis and interpretable neural architecture design.
📝 Abstract
We investigate the function-space optimality (specifically, the Banach-space optimality) of a large class of shallow neural architectures with multivariate nonlinearities/activation functions. To that end, we construct a new family of Banach spaces defined via a regularization operator, the $k$-plane transform, and a sparsity-promoting norm. We prove a representer theorem that states that the solution sets to learning problems posed over these Banach spaces are completely characterized by neural architectures with multivariate nonlinearities. These optimal architectures have skip connections and are tightly connected to orthogonal weight normalization and multi-index models, both of which have received recent interest in the neural network community. Our framework is compatible with a number of classical nonlinearities including the rectified linear unit (ReLU) activation function, the norm activation function, and the radial basis functions found in the theory of thin-plate/polyharmonic splines. We also show that the underlying spaces are special instances of reproducing kernel Banach spaces and variation spaces. Our results shed light on the regularity of functions learned by neural networks trained on data, particularly with multivariate nonlinearities, and provide new theoretical motivation for several architectural choices found in practice.