🤖 AI Summary
This work investigates whether single-hidden-layer multilayer perceptrons (MLPs), trained via standard optimization, spontaneously develop the non-smooth, high-dimensional nested functional geometry characterized by the Kolmogorov–Arnold (KA) superposition theorem.
Method: We introduce geometric proxy metrics based on exterior powers of the Jacobian matrix—specifically, the count of zero rows and the distribution of minors—to systematically quantify the dynamic emergence of KA structure during training.
Contribution/Results: We provide the first empirical evidence that KA geometry arises ubiquitously under conventional optimization, with its incidence exhibiting reproducible empirical dependencies on target function complexity, network width, and learning rate. Beyond revealing an intrinsic structured learning mechanism in shallow networks, our approach yields the first differentially geometric, monitorable criterion for KA emergence—offering a novel paradigm for understanding how neural networks implicitly construct task-beneficial representations.
📝 Abstract
The Kolmogorov-Arnold (KA) representation theorem constructs universal, but highly non-smooth inner functions (the first layer map) in a single (non-linear) hidden layer neural network. Such universal functions have a distinctive local geometry, a "texture," which can be characterized by the inner function's Jacobian $J({mathbf{x}})$, as $mathbf{x}$ varies over the data. It is natural to ask if this distinctive KA geometry emerges through conventional neural network optimization. We find that indeed KA geometry often is produced when training vanilla single hidden layer neural networks. We quantify KA geometry through the statistical properties of the exterior powers of $J(mathbf{x})$: number of zero rows and various observables for the minor statistics of $J(mathbf{x})$, which measure the scale and axis alignment of $J(mathbf{x})$. This leads to a rough understanding for where KA geometry occurs in the space of function complexity and model hyperparameters. The motivation is first to understand how neural networks organically learn to prepare input data for later downstream processing and, second, to learn enough about the emergence of KA geometry to accelerate learning through a timely intervention in network hyperparameters. This research is the "flip side" of KA-Networks (KANs). We do not engineer KA into the neural network, but rather watch KA emerge in shallow MLPs.