Generalized Convexity and Smoothness via Conjugate Duality: Optimization Theory for Deep Neural Networks

📅 2026-08-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Classical optimization theory struggles to explain the success of deep neural network (DNN) training due to its reliance on assumptions such as differentiability, convexity, or smoothness. This work addresses this gap by generalizing convexity and smoothness through Legendre functions and convex conjugates, introducing ℋ(ψ)-convexity and ℋ(Ψ)-smoothness, and revealing their duality. Building on this foundation, the authors develop a unified optimization framework that dispenses with traditional assumptions and propose a generalized gradient descent algorithm. They prove that this algorithm achieves optimality with a learning rate of 1 and establish rigorous convergence rate guarantees. By integrating composite optimization modeling, gradient energy analysis, and Jacobian-induced norm control, the theoretical predictions align closely with empirical training dynamics observed in experiments.
📝 Abstract
Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully explain this success. This limitation arises because conventional analyses rely on assumptions such as differentiability, convexity, or smoothness, which are often violated by DNN objectives. In this paper, we establish a unified optimization framework for DNN training by generalizing classical convexity and smoothness through Legendre functions and convex conjugation. Specifically, we introduce $\mathcal{H}(ψ)$-convexity and $\mathcal{H}(Ψ)$-smoothness, which unify convex and non-convex as well as smooth and non-smooth objectives within a single formalism and reveal a natural duality between generalized smoothness and convexity. Building on these generalized properties, we introduce generalized gradient descent (GD) and generalized SGD through convex conjugation. We theoretically prove that generalized GD admits an optimal learning rate of exactly $1$, and derive rigorous gradient-energy-based convergence rates for both proposed optimizers. We further reformulate DNN training as a composite optimization problem, demonstrating that its convergence relies on jointly reducing the gradient energy and controlling the induced norm of the network Jacobian. To characterize the practical influences of network architectures and training configurations, we introduce the gradient correlation factor and model capacity risk, and quantitatively analyze how architectural designs, batch size, and model capacity shape training convergence. Extensive experiments across diverse network architectures, datasets, optimizers, and loss functions validate our theoretical bounds and demonstrate precise alignment between our theoretical predictions and empirical training dynamics.
Problem

Research questions and friction points this paper is trying to address.

deep neural networks
optimization theory
non-convex optimization
non-smooth optimization
stochastic gradient descent
Innovation

Methods, ideas, or system contributions that make the work stand out.

generalized convexity
conjugate duality
gradient energy
non-convex optimization
deep neural network training
B
Binchuan Qi
College of Electronics and Information Engineering, Tongji University; Sihao Street, Zhejiang Yuying College of Vocational Technology