Score
Design and implement a neural network output head that parameterizes a convex potential with an input-convex neural network (ICNN) and uses that potential to compute Bregman divergences as scores. Ensure the module enforces convexity and local curvature properties of the potential so that computed Bregman divergences are nonnegative and suitable as a scorer or loss.
ReLU neural networks suffer from inaccurate convexity modeling in mathematical imaging tasks—such as optimization-based reconstruction and optimal transport—limiting their theoretical reliability and practical applicability. Method: We establish the first necessary and sufficient condition for convexity of arbitrary-depth ReLU networks, revealing that single-hidden-layer networks are equivalent to input-convex neural networks (ICNNs), whereas deeper ICNNs exhibit fundamental representational limitations. We develop a general convexity certification framework grounded in weight product analysis and activation pattern enumeration, integrating path-augmentation techniques, piecewise-affine modeling, and convex analysis. Furthermore, we design a scalable, exact convexity verification algorithm. Results: Our approach enables efficient convexity certification for large-scale piecewise-affine ReLU networks—achieving the first such capability—thereby transcending architectural constraints of ICNNs and providing both theoretical foundations and computational tools for trustworthy deployment of deep learning in convex optimization-driven imaging.
This work addresses the computational challenges of embedding neural networks into mathematical optimization, where conventional feedforward neural networks (FNNs) yield mixed-integer programming (MIP) reformulations that are computationally expensive and suffer from loose relaxations. To overcome these limitations, the paper proposes using input convex neural networks (ICNNs) as surrogate models, leveraging their inherent convexity to construct tight linear programming (LP) relaxations. The authors establish, for the first time, an exact convex hull-based continuous relaxation of ICNNs over box domains, yielding an LP representation free of integrality gaps. Furthermore, they introduce a novel branch-and-bound algorithm that branches directly on input variables. Demonstrated across applications in humanitarian food aid allocation, oil well trajectory planning, and wine blending, the approach achieves approximation accuracy comparable to FNNs while significantly improving solution speed and scalability.
To address function convex approximation—critical in optimal transport and related problems—this paper proposes Input-Convex Kolmogorov–Arnold Networks (ICKANs), the first neural architecture leveraging the Kolmogorov–Arnold representation theorem for input-convex network design. We construct two novel architectures: piecewise-linear ICKANs and differentiable cubic-spline ICKANs, both rigorously enforcing input convexity. Theoretically, we prove that piecewise-linear ICKANs possess universal approximation capability for convex functions, while cubic-spline ICKANs balance smoothness and expressive power. Empirically, ICKANs match the performance of classical Input-Convex Neural Networks (ICNNs) on convex function approximation and optimal transport benchmarks. Notably, cubic-spline ICKANs achieve accuracy on par with ICNNs in Wasserstein distance estimation, demonstrating their effectiveness, theoretical soundness, and practical utility.
This work systematically investigates the geometric and topological properties of the loss landscape of regularized neural networks, focusing on critical point structure, connectivity of global minima, existence of non-increasing-loss paths between optima, and non-uniqueness of global solutions—revealing a width-dependent topological phase transition. Using convex duality, we reformulate the optimization problem and rigorously characterize the structure of the critical point set and the global minimum set. We prove, for the first time, that any two global minima are connected by a continuous path along which the loss is everywhere non-increasing. We construct explicit counterexamples exhibiting a continuum of global minima, confirming the width-driven topological phase transition. These results extend to vector-valued outputs and parallel three-layer networks. Collectively, they establish a scalable, architecture-agnostic theory of global minimum connectivity and solution-set geometry, offering new insights into generalization and optimization in deep learning.
Practical implementation of Brenier’s polar decomposition theorem in machine learning—decomposing an arbitrary vector field into the gradient of a convex potential and a measure-preserving map—is hindered by the latter’s frequent non-injectivity, rendering its inverse unreliable. Method: We propose the first differentiable, learnable neural decomposition framework: (1) parameterize the convex potential via input-convex neural networks (ICNNs), compute its gradient using explicit convex conjugation and conjugate gradient optimization; (2) adopt a dual-path architecture—auxiliary networks estimate the measure-preserving map while stochastic generators mitigate ill-posedness in inversion; (3) enforce measure preservation through neural optimal transport theory. Contribution/Results: Experiments demonstrate significantly improved convergence in non-convex optimization and high-fidelity sampling from non-log-concave densities, establishing, for the first time, the practical feasibility of Brenier decomposition in high-dimensional, non-convex settings.
This work addresses the optimization challenges in Input Convex Neural Networks (ICNNs), where non-negative weight constraints often lead to vanishing gradients and training stagnation. To overcome these limitations, the authors propose a hypernetwork-based “lift” framework that generates ICNN weights from permutation-invariant summaries of input batches via an unconstrained hypernetwork. The approach incorporates learnable biases, batch conditioning, and a cross-covariance regularization term to soften the loss landscape and alleviate optimization plateaus. Evaluated on log-concave energy modeling and convex potential normalizing flows, the method significantly outperforms projection-based gradient descent and Softplus reparameterization, achieving lower test losses and enabling training trajectories to transition from flat plateaus to sustained descent.
Input Convex Neural Networks (ICNNs) are commonly used in a two-stage manner: one first trains a convex network and then minimizes it over its input in a downstream inference problem. Recent second-order-cone ICNNs (SOC-ICNNs) enrich ReLU-based ICNNs with quadratic and conic modules and admit an exact representation as value functions of second-order cone programs (SOCPs). This value-function structure enables an explicit convex-analytic treatment of SOC-ICNN inference. In this paper, we study the exact first-order and local second-order geometry of SOC-ICNNs from the dual viewpoint. We show that supporting slopes, subdifferentials, directional derivatives, and local Hessians can be recovered directly from optimal dual variables. These results provide the geometric primitives for white-box SOC-ICNN inference, going beyond black-box automatic differentiation. Numerical experiments validate the exact multiplier readout, the local Hessian formula, and the set-valued behavior at structurally degenerate inputs. We also provide a step-by-step tutorial showing how the readout mechanism instantiates a complete white-box inference loop. The code is available at https://anonymous.4open.science/r/SOC-ICNN-Theory-BEFC/.
This work addresses the limited parameter efficiency and poor scalability of existing Input Convex Neural Networks (ICNNs) in shape-constrained learning and high-dimensional optimal transport. To overcome these limitations, the authors propose Hyper Input Convex Neural Networks (HyCNNs), which integrate the Maxout activation mechanism into the ICNN architecture. This design strictly preserves input convexity while substantially enhancing model expressivity and scalability. Theoretical analysis demonstrates that HyCNNs require exponentially fewer parameters than ICNNs to approximate quadratic functions. Empirical evaluations on convex regression, interpolation, and high-dimensional optimal transport tasks using single-cell RNA sequencing data show that HyCNNs significantly outperform baseline methods, including standard ICNNs and multilayer perceptrons (MLPs).
This work addresses the challenges of parameterizing convex sets in shape optimization and inverse design by proposing an implicit representation based on sublinear neural networks. The method flexibly characterizes arbitrary convex bodies by learning positively homogeneous and convex support and gauge functions. It enjoys theoretical universal approximation capabilities for convex sets and demonstrates strong empirical performance, accurately reconstructing target shapes in experiments, thereby validating its expressiveness and effectiveness. The key innovation lies in integrating convex analysis with neural networks to establish a convex set parameterization framework that simultaneously offers rigorous theoretical guarantees and practical performance.
This study addresses the computational complexity of certifying global Lipschitz constants for input convex neural networks by integrating parameterized complexity theory, lifting selector reductions, and quantitative theorems on rational cyclic zonogons. We prove that this decision problem is NP-complete and W[1]-hard, establishing a theoretical barrier that precludes the separation of dimension and precision parameters. Consequently, our results rule out fixed-parameter tractability and dimension-independent approximation algorithms. This work resolves an open question from COLT 2025 regarding the Euclidean setting and demonstrates that convexity cannot circumvent the inherent computational bottlenecks in global sensitivity analysis. These findings fundamentally clarify the limits of efficient verification for this class of structured networks, highlighting persistent hardness despite architectural constraints.