🤖 AI Summary
To address the degradation of deep neural networks’ generalization under distributional shift, this paper proposes a novel regularizer that directly minimizes the spectral radius of the loss function’s Hessian matrix, thereby guiding optimization toward flat minima. It is the first work to incorporate the spectral radius as a differentiable regularizer within non-convex optimization, with theoretical convergence guarantees established. We design an efficient gradient approximation algorithm leveraging Hessian-vector products, implicit differentiation, and adaptive step sizing—enabling practical integration into standard training pipelines. Extensive experiments on real-world cross-distribution benchmarks—including medical imaging—demonstrate consistent and significant improvements in out-of-distribution generalization across multiple evaluation metrics, outperforming state-of-the-art baselines. The core contributions are: (i) a spectral-radius-driven, differentiable flatness regularizer; and (ii) its scalable, implementation-friendly instantiation for deep neural networks.
📝 Abstract
We develop a regularization method which finds flat minima during the training of deep neural networks and other machine learning models. These minima generalize better than sharp minima, allowing models to better generalize to real word test data, which may be distributed differently from the training data. Specifically, we propose a method of regularized optimization to reduce the spectral radius of the Hessian of the loss function. Additionally, we derive algorithms to efficiently perform this optimization on neural networks and prove convergence results for these algorithms. Furthermore, we demonstrate that our algorithm works effectively on multiple real world applications in multiple domains including healthcare. In order to show our models generalize well, we introduce different methods of testing generalizability.