🤖 AI Summary
This work systematically investigates the geometric and topological properties of the loss landscape of regularized neural networks, focusing on critical point structure, connectivity of global minima, existence of non-increasing-loss paths between optima, and non-uniqueness of global solutions—revealing a width-dependent topological phase transition. Using convex duality, we reformulate the optimization problem and rigorously characterize the structure of the critical point set and the global minimum set. We prove, for the first time, that any two global minima are connected by a continuous path along which the loss is everywhere non-increasing. We construct explicit counterexamples exhibiting a continuum of global minima, confirming the width-driven topological phase transition. These results extend to vector-valued outputs and parallel three-layer networks. Collectively, they establish a scalable, architecture-agnostic theory of global minimum connectivity and solution-set geometry, offering new insights into generalization and optimization in deep learning.
📝 Abstract
We discuss several aspects of the loss landscape of regularized neural networks: the structure of stationary points, connectivity of optimal solutions, path with nonincreasing loss to arbitrary global optimum, and the nonuniqueness of optimal solutions, by casting the problem into an equivalent convex problem and considering its dual. Starting from two-layer neural networks with scalar output, we first characterize the solution set of the convex problem using its dual and further characterize all stationary points. With the characterization, we show that the topology of the global optima goes through a phase transition as the width of the network changes, and construct counterexamples where the problem may have a continuum of optimal solutions. Finally, we show that the solution set characterization and connectivity results can be extended to different architectures, including two-layer vector-valued neural networks and parallel three-layer neural networks.