train monotonic icnns

Designs, implements, and trains input-convex neural network (ICNN) architectures that are monotonic in specified inputs, producing learned mappings that are convex and nondecreasing with respect to those inputs. This work includes enforcing monotonicity and convexity via architectural parameterizations (e.g., non‑negative weights), regularizers or loss penalties, input preprocessing or transforms, and procedures to verify and analyze constraint satisfaction during training.

trainmonotonicicnns

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Traditional monotonic neural networks rely on bounded activation functions and non-negative weight constraints, suffering from optimization difficulties and limited universal approximation capability. This work breaks this paradigm by proving, for the first time, that convex monotonic activations paired with non-positive weight constraints retain universal approximation power; it further establishes a theoretical equivalence between one-sided saturation in activations and sign constraints on weights. We propose a weight-sign–adaptive activation mechanism that eliminates the need for reparameterization. Theoretically, our framework expands the approximation capacity boundary of monotonic networks. Empirically, the proposed architecture demonstrates improved optimization stability, greater robustness to initialization, and significantly outperforms conventional monotonic MLPs across multiple benchmark tasks.

Establishing equivalence between activation saturation side and weight constraint signGeneralizing universal approximation for monotonic MLPs with non-negative weights and alternating saturation activationsProposing weight-sign-adaptive activations to ease optimization without reparameterization

Convexity in ReLU Neural Networks: beyond ICNNs?

Jan 06, 2025
AG
Anne Gagneux
🏛️ ENS de Lyon | Université Claude Bernard Lyon 1 | Inria | Université de Toulouse

ReLU neural networks suffer from inaccurate convexity modeling in mathematical imaging tasks—such as optimization-based reconstruction and optimal transport—limiting their theoretical reliability and practical applicability. Method: We establish the first necessary and sufficient condition for convexity of arbitrary-depth ReLU networks, revealing that single-hidden-layer networks are equivalent to input-convex neural networks (ICNNs), whereas deeper ICNNs exhibit fundamental representational limitations. We develop a general convexity certification framework grounded in weight product analysis and activation pattern enumeration, integrating path-augmentation techniques, piecewise-affine modeling, and convex analysis. Furthermore, we design a scalable, exact convexity verification algorithm. Results: Our approach enables efficient convexity certification for large-scale piecewise-affine ReLU networks—achieving the first such capability—thereby transcending architectural constraints of ICNNs and providing both theoretical foundations and computational tools for trustworthy deployment of deep learning in convex optimization-driven imaging.

convexity verificationoptimization algorithmsReLU neural networks

This work addresses the challenge of efficiently enforcing global hard monotonicity constraints in neural networks while preserving interpretability and training simplicity. To this end, the authors propose MKAN, a novel architecture based on Kolmogorov–Arnold Networks, which guarantees hard monotonicity across all parameters through exponential reparameterization of B-spline coefficients, positive edge weights, and monotonic basis activation functions, while remaining compatible with standard unconstrained gradient descent optimization. The key contributions include the first realization of full-parameter hard monotonicity within the KAN framework and an architecture-agnostic representation cost theorem that informs the design of monotonic encoder widths. Experiments demonstrate that MKAN achieves state-of-the-art performance on the SMM/ICML-2024 benchmark, validates the predicted 2N* width scaling law across four real-world datasets, and significantly outperforms KAN, MLP, and linear models in Spearman correlation alignment on controllable generation tasks.

functional transparencyinductive biasKolmogorov-Arnold Networks

Convex Formulations for Training Two-Layer ReLU Neural Networks

Oct 29, 2024
KP
Karthik Prakhya
🏛️ Ume˚a University | Imperial College

Training two-layer ReLU neural networks is inherently non-convex, posing significant theoretical and computational challenges. Method: This work establishes, for the first time in the infinite-width limit, an exact equivalence between ReLU network training and a finite-dimensional convex completely positive program (CPP). We propose a compact semidefinite programming (SDP) relaxation that is solvable in polynomial time and preserves the optimal value of the original CPP problem exactly. Contribution/Results: Theoretically, we derive a precise correspondence between non-convex neural network training and convex optimization. Empirically, our SDP-based approach achieves competitive test accuracy on multi-class classification benchmarks, empirically validating both the tightness of the relaxation and its generalization capability. This work provides a novel convex analytical framework and a tractable computational pathway for deep learning training, bridging classical convex optimization theory with modern neural network practice.

Evaluates relaxation tightness and performance in classification tasks.Introduces semidefinite relaxation for polynomial-time solvable convex formulations.Reformulates training infinite-width two-layer ReLU networks as convex optimization.

MonoKAN: Certified Monotonic Kolmogorov-Arnold Network

Sep 17, 2024
AP
Alejandro Polo-Molina
🏛️ Institute for Research in Technology (IIT) | Universidad Pontificia Comillas | Department of Applied Mathematics | Department of Quantitative Methods | ICADE | ICAI School of Engineering

Artificial neural networks (ANNs) suffer from poor interpretability and struggle to simultaneously satisfy expert-defined partial monotonicity constraints and achieve high predictive performance. Method: We propose the first trustworthy AI framework that integrates certified partial monotonicity with Kolmogorov–Arnold Networks (KANs). Our approach enforces monotonicity structurally—both in network topology and activation functions—via non-negative linear weights and learnable, strictly monotonic activation functions based on cubic Hermite splines. This ensures monotonicity is provable, interpretable, and differentiable end-to-end. Contribution/Results: The framework achieves certified partial monotonicity without compromising expressivity or trainability. Extensive experiments on multiple benchmark tasks demonstrate substantial improvements over state-of-the-art monotonic MLPs, delivering both higher prediction accuracy and rigorous, verifiable monotonicity guarantees—thereby advancing reliable, domain-aligned AI deployment.

Achieving certified partial monotonicity in neural networksEnhancing model interpretability while maintaining performanceEnsuring input-output relationships follow expert-imposed constraints

Latest Papers

What's happening recently
View more

This work addresses the computational challenges of embedding neural networks into mathematical optimization, where conventional feedforward neural networks (FNNs) yield mixed-integer programming (MIP) reformulations that are computationally expensive and suffer from loose relaxations. To overcome these limitations, the paper proposes using input convex neural networks (ICNNs) as surrogate models, leveraging their inherent convexity to construct tight linear programming (LP) relaxations. The authors establish, for the first time, an exact convex hull-based continuous relaxation of ICNNs over box domains, yielding an LP representation free of integrality gaps. Furthermore, they introduce a novel branch-and-bound algorithm that branches directly on input variables. Demonstrated across applications in humanitarian food aid allocation, oil well trajectory planning, and wine blending, the approach achieves approximation accuracy comparable to FNNs while significantly improving solution speed and scalability.

convexityinput convex neural networksmathematical optimization

This work addresses the optimization challenges in Input Convex Neural Networks (ICNNs), where non-negative weight constraints often lead to vanishing gradients and training stagnation. To overcome these limitations, the authors propose a hypernetwork-based “lift” framework that generates ICNN weights from permutation-invariant summaries of input batches via an unconstrained hypernetwork. The approach incorporates learnable biases, batch conditioning, and a cross-covariance regularization term to soften the loss landscape and alleviate optimization plateaus. Evaluated on log-concave energy modeling and convex potential normalizing flows, the method significantly outperforms projection-based gradient descent and Softplus reparameterization, achieving lower test losses and enabling training trajectories to transition from flat plateaus to sustained descent.

gradient attenuationinput-convex neural networksnon-negative weights

Input Convex Neural Networks (ICNNs) are commonly used in a two-stage manner: one first trains a convex network and then minimizes it over its input in a downstream inference problem. Recent second-order-cone ICNNs (SOC-ICNNs) enrich ReLU-based ICNNs with quadratic and conic modules and admit an exact representation as value functions of second-order cone programs (SOCPs). This value-function structure enables an explicit convex-analytic treatment of SOC-ICNN inference. In this paper, we study the exact first-order and local second-order geometry of SOC-ICNNs from the dual viewpoint. We show that supporting slopes, subdifferentials, directional derivatives, and local Hessians can be recovered directly from optimal dual variables. These results provide the geometric primitives for white-box SOC-ICNN inference, going beyond black-box automatic differentiation. Numerical experiments validate the exact multiplier readout, the local Hessian formula, and the set-valued behavior at structurally degenerate inputs. We also provide a step-by-step tutorial showing how the readout mechanism instantiates a complete white-box inference loop. The code is available at https://anonymous.4open.science/r/SOC-ICNN-Theory-BEFC/.

convex neural networksdual geometrysecond-order cone program

This study addresses the computational complexity of certifying global Lipschitz constants for input convex neural networks by integrating parameterized complexity theory, lifting selector reductions, and quantitative theorems on rational cyclic zonogons. We prove that this decision problem is NP-complete and W[1]-hard, establishing a theoretical barrier that precludes the separation of dimension and precision parameters. Consequently, our results rule out fixed-parameter tractability and dimension-independent approximation algorithms. This work resolves an open question from COLT 2025 regarding the Euclidean setting and demonstrates that convexity cannot circumvent the inherent computational bottlenecks in global sensitivity analysis. These findings fundamentally clarify the limits of efficient verification for this class of structured networks, highlighting persistent hardness despite architectural constraints.

Approximation barriersInput-convex neural networksLipschitz constant

Hot Scholars

BC

Bing Chen

PhD Student, University of Waterloo
machine learningneural networks
JT

Julián Tachella

CNRS research scientist at ENS de Lyon
Signal ProcessingImage ProcessingMachine Learning
RB

Ray Bai

Department of Statistics, George Mason University
statisticsmachine learningBayesian statisticsdeep learning
LW

Lin Wang

University of Jinan, 250022
Machine LearningData MiningScientific Computing