🤖 AI Summary
This work addresses the limitations of conventional Gaussian-noise perturbations by systematically uncovering previously overlooked one-dimensional (1D) and two-dimensional (2D) local geometric structures in deep neural network (DNN) loss landscapes. Methodologically, we introduce a progressive taxonomy of five types of 1D loss curves (e.g., *v-basin*, *vvv-basin*) and design a perturbation-direction mining algorithm that integrates low-dimensional subspace projection with Hessian spectral analysis to automatically extract and visualize complex geometric structures. Our contributions include: (i) the first empirical observation and visualization of canonical 2D loss structures—including saddle surfaces and “bottle-bottom” geometries—in real DNNs; (ii) a theoretical characterization linking the geometric properties of perturbation directions to the eigenvalue distribution of the Hessian; and (iii) a novel geometric perspective for understanding generalization behavior and optimization dynamics.
📝 Abstract
The loss landscape of deep neural networks (DNNs) is commonly considered complex and wildly fluctuated. However, an interesting observation is that the loss surfaces plotted along Gaussian noise directions are almost v-basin ones with the perturbed model lying on the basin. This motivates us to rethink whether the 1D or 2D subspace could cover more complex local geometry structures, and how to mine the corresponding perturbation directions. This paper systematically and gradually categorizes the 1D curves from simple to complex, including v-basin, v-side, w-basin, w-peak, and vvv-basin curves. Notably, the latter two types are already hard to obtain via the intuitive construction of specific perturbation directions, and we need to propose proper mining algorithms to plot the corresponding 1D curves. Combining these 1D directions, various types of 2D surfaces are visualized such as the saddle surfaces and the bottom of a bottle of wine that are only shown by demo functions in previous works. Finally, we propose theoretical insights from the lens of the Hessian matrix to explain the observed several interesting phenomena.