🤖 AI Summary
This work investigates the algorithmic learnability gap between constant-depth and logarithmic-depth neural networks, addressing a key limitation in existing depth-separation theory—namely, the absence of efficient learning algorithms that realize such separations. By constructing a class of Boolean functions with hierarchical Fourier spectra and leveraging tools from Fourier analysis, spectral norm control, and layerwise coordinate descent optimization, the paper establishes the first algorithmic separation in practical learnability: logarithmic-depth networks can efficiently learn this function class, whereas any constant-depth network of polynomial width incurs a constant L² approximation error under the uniform distribution over the hypercube.
📝 Abstract
Despite the empirical advantages of deep networks over shallow ones, theoretical depth separations largely concern approximation power, while algorithmic results are mostly limited to comparisons between two- and three-layer networks. In this work, we prove the first algorithmic separation between constant-depth and logarithmic-depth networks.
Specifically, we identify a class of Boolean functions with hierarchically structured Fourier spectra that logarithmic-depth networks can learn efficiently using layerwise coordinate descent by reconstructing the spectra hierarchically and adaptively. We also exhibit a subclass for which every constant-depth, polynomial-width network with sufficiently regular activations and controlled spectral norms must incur constant $L^2$ approximation error under the uniform distribution over the hypercube.