Beyond Overparameterization: Provable Learning of Input-Convex Multi-Layer Polynomial Networks with Active Queries

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses longstanding theoretical limitations in analyzing deep neural networks, specifically the reliance on overparameterization assumptions, high sample complexity, and the absence of parameter-level recovery guarantees. To overcome these challenges, this work proposes the ASPIRE algorithm for input-convex polynomial networks, which integrates active querying strategies, sampling-based diagonalization techniques, and iterative feature direction extraction to achieve precise layer-wise sampling. The primary contribution is the first exact parameter recovery guarantee under conditions where network depth grows exponentially with a polynomial. Furthermore, it is proven that all parameters can be recovered to δ-precision in polynomial time, substantially reducing sample complexity and establishing new theoretical bounds. These results rigorously validate the critical role of high-quality data in enabling efficient network training.
📝 Abstract
The theoretical understanding of multi-layer neural networks is largely confined to overparameterized settings, which obscure parameter identifiability and incur high sample complexity. Neural tangent kernel (NTK) provides a general theory for wide networks, but does not offer efficient sample-complexity guarantees. Recent feature-learning results go beyond kernel methods for single-neuron, multi-index, and hierarchical targets. However, the analysis is often restricted to shallow or specific architectures and to the overparameterized regime. We break this paradigm to achieve parameter-level recovery of deep target networks, albeit by using active data queries. Specifically, we study $L$-layer polynomial networks with even degree-$k$ monomial activations and nonnegative higher-layer weights. This structure makes the target network input-convex, while the optimization landscape remains highly nonconvex with respect to the parameters. Leveraging input convexity and active queries, we propose \textbf{ASPIRE} (\textbf{A}ctive \textbf{S}am\textbf{P}ling for \textbf{I}terative \textbf{R}ecovery via \textbf{E}igendirections), a layerwise sampling-based diagonalization algorithm that recovers all network parameters to $\delta$-accuracy with sample complexity $ \widetilde O_{k,L}\left(d^{L^2+O(L)}\delta^{-2e}\right) $ in polynomial time. To our knowledge, this is the \emph{first} parameter-recovery guarantee for deep target networks whose exponent grows only polynomially with depth, as well as the \emph{first} justification for the effectiveness of using high-quality data in neural network training, with a remarkably \emph{exponential} separation.
Problem

Research questions and friction points this paper is trying to address.

overparameterization
parameter recovery
sample complexity
deep neural networks
input-convex networks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Input-Convex Networks
Active Queries
Parameter Recovery
Polynomial Networks
ASPIRE Algorithm
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jinqi Tang
Peking University
Qian Chen
Qian Chen
Peking University
S
Shihong Ding
Peking University
Cong Fang
Cong Fang
Peking University
machine learningoptmizationstatistics