Score
Design and implement regularization methods that explicitly operate on the spectral (frequency) content of model weights, activations, or learned representations. This includes constructing loss terms, convolutional filters, or training procedures (e.g., Gaussian convolution of selected layers, frequency-band penalties) that attenuate low-pass bias and encourage sensitivity to mid-to-high frequency components.
The physical mechanisms by which L2 regularization and dropout implicitly impose low-frequency inductive biases in CNNs remain poorly understood. Method: We propose a visual diagnostic framework and introduce the Spectral Suppression Ratio (SSR) to quantify the low-pass filtering strength of regularizers. Using discrete radial spectral analysis, weight spectral evolution tracking, and experiments on ResNet-18/CIFAR-10, we systematically characterize how regularization shapes weight frequency spectra. Contribution/Results: We find L2 regularization reduces high-frequency weight energy by over 3×, significantly improving blur robustness (+6.2%) but degrading Gaussian noise robustness—revealing an intrinsic accuracy–robustness trade-off governed by spectral bias. Furthermore, we design a small-kernel anti-aliasing model that provides interpretable theoretical grounding for understanding spectral inductive biases in deep learning.
This work investigates the implicit spectral bias regulation mechanism of gradient descent (GD) in one-dimensional shallow neural networks. Theoretically, GD is shown to act as a singular-value shrinkage operator on the network’s Jacobian matrix, where the learning rate and iteration count jointly determine the spectral bandwidth—the number of active frequency components. Three key contributions are made: (1) GD is formally modeled as an explicit frequency-domain shrinkage operation, with a quantitative relationship established between hyperparameters and bandwidth; (2) it is proven that GD’s implicit regularization effect on spectral bias is valid only under monotonic activation functions; (3) non-monotonic activations—including sinc and Gaussian—are proposed as efficient alternatives, significantly enhancing spectral control efficiency. Collectively, these results provide a novel theoretical framework for understanding implicit spectral bias in deep learning.
This paper investigates the generalization performance of random feature methods under generalized spectral regularization—including explicit schemes (e.g., Tikhonov regularization) and implicit schemes (e.g., gradient descent and accelerated algorithms). Leveraging insights from neural tangent kernel (NTK) theory, it establishes optimal learning rates for the first time under non-RKHS-type regularity assumptions, thereby unifying the analysis of implicit and explicit regularization mechanisms. Methodologically, the work integrates random feature mappings, spectral analysis, source condition modeling, and rigorous generalization error bound derivation to obtain tight convergence rates for both neural networks and neural operators. Key contributions are: (1) extending optimal convergence guarantees beyond the RKHS framework to broader classes of spectral regularity; (2) providing a unified, tight, and improved theoretical bound applicable to diverse kernel-based algorithms; and (3) substantially enhancing computational efficiency and depth of generalization analysis for large-scale kernel methods.
To address the degradation of deep neural networks’ generalization under distributional shift, this paper proposes a novel regularizer that directly minimizes the spectral radius of the loss function’s Hessian matrix, thereby guiding optimization toward flat minima. It is the first work to incorporate the spectral radius as a differentiable regularizer within non-convex optimization, with theoretical convergence guarantees established. We design an efficient gradient approximation algorithm leveraging Hessian-vector products, implicit differentiation, and adaptive step sizing—enabling practical integration into standard training pipelines. Extensive experiments on real-world cross-distribution benchmarks—including medical imaging—demonstrate consistent and significant improvements in out-of-distribution generalization across multiple evaluation metrics, outperforming state-of-the-art baselines. The core contributions are: (i) a spectral-radius-driven, differentiable flatness regularizer; and (ii) its scalable, implementation-friendly instantiation for deep neural networks.
Kernel ridge regression (KRR) and Gaussian processes (GPs) suffer from degraded learnability on real-world data due to rapid eigenvalue decay of the kernel matrix, impairing their ability to recover target functions. Method: We propose a spectral learnability analysis framework that introduces the idealized data measure’s eigen-spectrum as a theoretical benchmark, quantifying spectral deviation of real data from this benchmark. Leveraging symmetry properties of the kernel operator, we derive tight theoretical bounds on learnability, revealing bottlenecks in learning components aligned with high-eigenvalue eigenvectors. The framework integrates kernel spectral analysis, eigenvalue estimation, and symmetry-driven generalization theory to quantitatively characterize the learnability of kernel methods under realistic data distributions. Results: Experiments demonstrate substantial improvements in predictive accuracy for learnability estimation. The framework establishes a novel paradigm for theoretically interpreting and practically deploying kernel methods on non-ideal, real-world data.
Existing diffusion models typically employ pointwise reconstruction losses that are insensitive to signal spectra and multiscale structures, often yielding samples with imbalanced frequency content and insufficient structural detail. To address this limitation, this work proposes a lightweight, plug-and-play spectral regularization framework that introduces differentiable Fourier- and wavelet-domain losses as soft inductive biases during standard training. The approach requires no modifications to the diffusion process, model architecture, or sampling procedure, and is compatible with mainstream paradigms such as DDPM, DDIM, and EDM. Evaluated on high-resolution unconditional generation tasks for both images and audio, the method consistently enhances sample quality—particularly in terms of spectral fidelity and multiscale coherence—while incurring negligible computational overhead.
This work addresses the instability of reconstruction in ill-posed inverse problems caused by noise, as well as limitations of conventional regularization methods—such as reliance on manual hyperparameter tuning—and the lack of interpretability and cross-resolution generalization in deep learning approaches. To this end, we propose the Spectral Correction Network (SC-Net), which learns a signal-to-noise-ratio-adaptive, pointwise filtering function in the spectral domain of the forward operator to reweight spectral coefficients, yielding stable and interpretable solutions. Our method uniquely integrates interpretable adaptive spectral filtering with operator learning, and we theoretically prove that it can approximate continuous inverse operators, possesses discretization invariance, and achieves minimax optimal convergence rates. Experiments on 1D integral equations demonstrate a convergence rate of $O(\delta^{0.5})$, matching theoretical optimality; the learned filter outperforms Oracle Tikhonov regularization and generalizes zero-shot from $N=256$ to $N=2048$ with reconstruction error maintained at approximately 0.23.
This study investigates the stability of spectral methods for coefficient estimation in supervised regression under additive label noise. Focusing on multi-basis sparse spectral representations, the authors derive a closed-form expression for the overlap between noisy and noise-free coefficient vectors via eigen-geometric whitening. Their analysis reveals that the degradation in spectral learning performance is governed by a single intrinsic noise scale and establishes a critical noise threshold beyond which the underlying function structure becomes irrecoverable. The theoretical framework applies to various orthogonal bases—including Fourier, Legendre, Bessel, and Haar—and is corroborated by numerical experiments demonstrating that spectral coefficient estimates become significantly unstable once the noise level exceeds this threshold, thereby impeding accurate recovery of the true function.
This study investigates the stability of implicit low-rank regularization in deep matrix factorization under noisy perturbations. By integrating spectral analysis, gradient flow dynamics, and perturbation theory, it characterizes how the spectral structure of the target matrix, initialization, and step size influence the existence of a low-rank phase. The work establishes, for the first time, explicit spectral conditions under which implicit low-rank regularization persists in the presence of perturbations. Theoretically, it proves that the optimization algorithm still converges to a low-rank solution despite noise, and derives explicit bounds on both iteration complexity and eigenvalue recovery error. These results quantitatively reveal how perturbation magnitude affects stability. Numerical experiments corroborate the theoretical findings.
This work addresses the theoretical gap in characterizing learning behavior and benign overfitting in the under-regularized regime of high-dimensional spectral algorithms. Under a proportional asymptotic setting where sample size and dimensionality grow at the same rate, the authors develop a unified analytical framework for a broad class of kernels satisfying spectral scaling and hypercontractivity conditions, leveraging high-dimensional asymptotics, spectral decomposition of dot-product kernels on the sphere, and source condition modeling. They provide the first complete characterization of the three-phase learning curve—over-regularized, under-regularized, and interpolation regimes—across the entire regularization path, precisely describing the asymptotic excess risk under various source conditions. The analysis reveals that benign overfitting is pervasive in both under-regularized and interpolation regimes, elucidating its underlying mechanism, and extends to kernel ridge regression in high dimensions over ℝᵈ, while also establishing a theoretical link between kernel methods and sequence models.