🤖 AI Summary
The physical mechanisms by which L2 regularization and dropout implicitly impose low-frequency inductive biases in CNNs remain poorly understood. Method: We propose a visual diagnostic framework and introduce the Spectral Suppression Ratio (SSR) to quantify the low-pass filtering strength of regularizers. Using discrete radial spectral analysis, weight spectral evolution tracking, and experiments on ResNet-18/CIFAR-10, we systematically characterize how regularization shapes weight frequency spectra. Contribution/Results: We find L2 regularization reduces high-frequency weight energy by over 3×, significantly improving blur robustness (+6.2%) but degrading Gaussian noise robustness—revealing an intrinsic accuracy–robustness trade-off governed by spectral bias. Furthermore, we design a small-kernel anti-aliasing model that provides interpretable theoretical grounding for understanding spectral inductive biases in deep learning.
📝 Abstract
Regularization techniques such as L2 regularization (Weight Decay) and Dropout are fundamental to training deep neural networks, yet their underlying physical mechanisms regarding feature frequency selection remain poorly understood. In this work, we investigate the Spectral Bias of modern Convolutional Neural Networks (CNNs). We introduce a Visual Diagnostic Framework to track the dynamic evolution of weight frequencies during training and propose a novel metric, the Spectral Suppression Ratio (SSR), to quantify the "low-pass filtering" intensity of different regularizers. By addressing the aliasing issue in small kernels (e.g., 3x3) through discrete radial profiling, our empirical results on ResNet-18 and CIFAR-10 demonstrate that L2 regularization suppresses high-frequency energy accumulation by over 3x compared to unregularized baselines. Furthermore, we reveal a critical Accuracy-Robustness Trade-off: while L2 models are sensitive to broadband Gaussian noise due to over-specialization in low frequencies, they exhibit superior robustness against high-frequency information loss (e.g., low resolution), outperforming baselines by >6% in blurred scenarios. This work provides a signal-processing perspective on generalization, confirming that regularization enforces a strong spectral inductive bias towards low-frequency structures.