🤖 AI Summary
This work addresses the non-convex optimization challenge in training two-layer ReLU networks by proposing the first *strictly equivalent convex reformulation*. Methodologically, it establishes—under zero regularization—that the original non-convex problem is *exactly equivalent* to a convex gated ReLU optimization problem. It further derives data-dependent approximation bounds and constructs a theoretical framework grounded in conic decomposition and localized convex model classes. To enhance tractability and generalization, the approach introduces polyhedral conic constraints and group ℓ₁ regularization. Optimization is performed via an accelerated proximal gradient method coupled with an augmented Lagrangian solver, yielding substantial computational gains. Experiments on MNIST and CIFAR-10 demonstrate scalability and superior speed over SGD and commercial interior-point solvers, while empirically validating both the theoretical equivalence and the efficacy of the regularization path.
📝 Abstract
We develop fast algorithms and robust software for convex optimization of two-layer neural networks with ReLU activation functions. Our work leverages a convex reformulation of the standard weight-decay penalized training problem as a set of group-$ell_1$-regularized data-local models, where locality is enforced by polyhedral cone constraints. In the special case of zero-regularization, we show that this problem is exactly equivalent to unconstrained optimization of a convex"gated ReLU"network with non-singular gates. For problems with non-zero regularization, we show that convex gated ReLU models obtain data-dependent approximation bounds for the ReLU training problem. To optimize the convex reformulations, we develop an accelerated proximal gradient method and a practical augmented Lagrangian solver. We show that these approaches are faster than standard training heuristics for the non-convex problem, such as SGD, and outperform commercial interior-point solvers. Experimentally, we verify our theoretical results, explore the group-$ell_1$ regularization path, and scale convex optimization for neural networks to image classification on MNIST and CIFAR-10.