Optimal Tradeoffs Between Network Size and Parameter Magnitude in Neural Approximation and Minimax Regression

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文研究了神经网络大小与参数幅度之间的最优权衡,通过使用特定激活函数和最小二乘法,在保持统计准确性的前提下优化了网络结构。
📝 Abstract
The statistical accuracy of neural networks depends on both their approximation power and the complexity of the class fitted from data. While increasing network size is a natural way to improve approximation, parameter magnitude provides another resource whose role must be quantified in both respects. We establish a sharp width--magnitude tradeoff at fixed depth using one elementary bounded $1$-Lipschitz Dyadic--Triangular Activation. For the unit $β$-Hölder ball on $[0,1]^d$ with $0<β\leq1$, the optimal $L^p$ approximation error for $0<p<\infty$ is of order $[N^2\log(eNT)]^{-β/d}$ when the network width satisfies $N\geq2d+3$ and the parameter magnitudes are bounded by $T\geq1$. Matching lower bounds hold for every fixed globally Hölder activation; its Hölder exponent affects the constants but not the rate. Under bounded design densities and independent centered sub-Gaussian noise, approximate least squares over the full clipped class at depth $23$ attains the classical Hölder minimax risk $\mathcal{O}(M^{-\frac{2β}{2β+d}})$ without logarithmic loss whenever $N^2\log(eNT)\asymp M^{\frac{d}{2β+d}}$, where $M$ is the sample size. This yields a continuum of statistically optimal choices, ranging from unit parameter radius to fixed network size. At fixed size, four hidden layers with at most $8d+7$ nonzero parameters give a near-optimal radius, while six layers with at most $8d+27$ attain the optimal order $\log T=\mathcal{O}(η^{-d/β})$ at approximation error $η$. The same decoding method also yields fixed-size Transformer approximation.
Problem

Research questions and friction points this paper is trying to address.

network size
parameter magnitude
neural approximation
minimax regression
statistical accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

width-magnitude tradeoff
Hölder minimax risk
network depth
approximation error
parameter magnitude
🔎 Similar Papers
No similar papers found.