Score
Designs, implements, and trains parameterized potential (energy) functions—including unnormalized densities represented by input-convex neural networks or folded-normalization parameterizations—so that the learned mapping is convex in the appropriate arguments. Constructs strictly convex training objectives (e.g., by folding normalization into the loss), develops the corresponding optimization procedures, and analyzes convergence and stability properties of empirical minimizers in finite- and high-dimensional settings.
Training two-layer ReLU neural networks is inherently non-convex, posing significant theoretical and computational challenges. Method: This work establishes, for the first time in the infinite-width limit, an exact equivalence between ReLU network training and a finite-dimensional convex completely positive program (CPP). We propose a compact semidefinite programming (SDP) relaxation that is solvable in polynomial time and preserves the optimal value of the original CPP problem exactly. Contribution/Results: Theoretically, we derive a precise correspondence between non-convex neural network training and convex optimization. Empirically, our SDP-based approach achieves competitive test accuracy on multi-class classification benchmarks, empirically validating both the tightness of the relaxation and its generalization capability. This work provides a novel convex analytical framework and a tractable computational pathway for deep learning training, bridging classical convex optimization theory with modern neural network practice.
This paper addresses non-convex optimization in shallow neural networks across three fundamental tasks: exact representation, function approximation, and regression. We propose a unified convexification framework grounded in mean-field theory. Theoretically, we rigorously prove that the convexified problem admits no relaxation gap and derive an interpretable, closed-form generalization bound that explicitly characterizes hyperparameter influence and provides principled guidelines for optimal selection. Algorithmically, we design a task-adaptive solver: for low-dimensional settings, we employ the simplex method with theoretical guarantees of exact recovery; for high-dimensional settings, we combine sparsification with gradient descent to achieve efficient approximation. Empirical results demonstrate substantial improvements in test performance over standard training heuristics. To our knowledge, this is the first work achieving unified modeling, gap-free convexification, and joint optimization of generalization and algorithmic efficiency across all three tasks.
Training two-layer ReLU networks faces challenges including non-convex optimization prone to local minima, sensitivity to initialization and hyperparameters, and existing convexification methods suffering from prohibitive computational complexity—exponential or cubic in the number of neurons. Method: We propose the first efficient algorithm that simultaneously guarantees global convergence and achieves quadratic time complexity. Our approach integrates a dual-path strategy combining the Alternating Direction Method of Multipliers (ADMM) with sampling-based convex programming to construct a scalable approximate convex solver. Theoretically, we establish the first provably robust convex adversarial training framework. Results: Experiments demonstrate linear global convergence, with high accuracy attained already in the first iteration. The method significantly improves generalization and robustness under both standard and adversarial training settings, while closely approximating the global optimum.
This work addresses the non-convex optimization challenge in training two-layer ReLU networks by proposing the first *strictly equivalent convex reformulation*. Methodologically, it establishes—under zero regularization—that the original non-convex problem is *exactly equivalent* to a convex gated ReLU optimization problem. It further derives data-dependent approximation bounds and constructs a theoretical framework grounded in conic decomposition and localized convex model classes. To enhance tractability and generalization, the approach introduces polyhedral conic constraints and group ℓ₁ regularization. Optimization is performed via an accelerated proximal gradient method coupled with an augmented Lagrangian solver, yielding substantial computational gains. Experiments on MNIST and CIFAR-10 demonstrate scalability and superior speed over SGD and commercial interior-point solvers, while empirically validating both the theoretical equivalence and the efficacy of the regularization path.
This work systematically investigates the geometric and topological properties of the loss landscape of regularized neural networks, focusing on critical point structure, connectivity of global minima, existence of non-increasing-loss paths between optima, and non-uniqueness of global solutions—revealing a width-dependent topological phase transition. Using convex duality, we reformulate the optimization problem and rigorously characterize the structure of the critical point set and the global minimum set. We prove, for the first time, that any two global minima are connected by a continuous path along which the loss is everywhere non-increasing. We construct explicit counterexamples exhibiting a continuum of global minima, confirming the width-driven topological phase transition. These results extend to vector-valued outputs and parallel three-layer networks. Collectively, they establish a scalable, architecture-agnostic theory of global minimum connectivity and solution-set geometry, offering new insights into generalization and optimization in deep learning.
This work addresses the challenges of parameterizing convex sets in shape optimization and inverse design by proposing an implicit representation based on sublinear neural networks. The method flexibly characterizes arbitrary convex bodies by learning positively homogeneous and convex support and gauge functions. It enjoys theoretical universal approximation capabilities for convex sets and demonstrates strong empirical performance, accurately reconstructing target shapes in experiments, thereby validating its expressiveness and effectiveness. The key innovation lies in integrating convex analysis with neural networks to establish a convex set parameterization framework that simultaneously offers rigorous theoretical guarantees and practical performance.
This work addresses the challenges of non-convex optimization and theoretical gaps in shallow neural network training by reframing the discrete parameter learning problem as a continuous variational problem over parameter densities in a weighted Sobolev space. The authors propose a globally well-posed and stable optimization framework grounded in λ-convex functionals, innovatively integrating elliptic regularity theory with variational analysis to bypass conventional iterative optimization. Instead, the optimal parameter density is obtained directly by solving a single linear system, thereby unifying the neural tangent kernel (NTK) and feature learning perspectives. Theoretically, the solution converges to the continuous optimum at a rate of O(1/N), and the generalization error is bounded by O(1/α), significantly enhancing both training efficiency and theoretical interpretability.
Classical optimization theory struggles to explain the success of deep neural network (DNN) training due to its reliance on assumptions such as differentiability, convexity, or smoothness. This work addresses this gap by generalizing convexity and smoothness through Legendre functions and convex conjugates, introducing ℋ(ψ)-convexity and ℋ(Ψ)-smoothness, and revealing their duality. Building on this foundation, the authors develop a unified optimization framework that dispenses with traditional assumptions and propose a generalized gradient descent algorithm. They prove that this algorithm achieves optimality with a learning rate of 1 and establish rigorous convergence rate guarantees. By integrating composite optimization modeling, gradient energy analysis, and Jacobian-induced norm control, the theoretical predictions align closely with empirical training dynamics observed in experiments.
This work addresses the challenge of optimizing complex loss functions—such as Kullback–Leibler divergence or PDE residuals—over nonlinear manifolds like neural or tensor networks, where conventional and natural gradient descent often converge to suboptimal local minima with inefficient update directions. The paper introduces, for the first time, a momentum-augmented variant of natural gradient descent that integrates Heavy-Ball and Nesterov-type inertial mechanisms into a manifold-aware optimization framework. By leveraging tangent-space projections and Gram matrix preconditioning, the method achieves momentum-driven, locally optimal updates directly in function space. Empirical results demonstrate substantial improvements in convergence behavior for tasks including density estimation and physics-informed learning, effectively mitigating poor local minima and enhancing the quality of optimization trajectories.
This work addresses the optimization challenges in Input Convex Neural Networks (ICNNs), where non-negative weight constraints often lead to vanishing gradients and training stagnation. To overcome these limitations, the authors propose a hypernetwork-based “lift” framework that generates ICNN weights from permutation-invariant summaries of input batches via an unconstrained hypernetwork. The approach incorporates learnable biases, batch conditioning, and a cross-covariance regularization term to soften the loss landscape and alleviate optimization plateaus. Evaluated on log-concave energy modeling and convex potential normalizing flows, the method significantly outperforms projection-based gradient descent and Softplus reparameterization, achieving lower test losses and enabling training trajectories to transition from flat plateaus to sustained descent.