New Complexity-Theoretic Frontiers of Tractability for Neural Network Training

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
The computational complexity of optimally training neural networks has long remained unclear, particularly due to the lack of a systematic characterization of architectures solvable in polynomial time. This work addresses this gap through complexity-theoretic analysis and algorithm design, establishing for the first time that ReLU networks admit polynomial-time optimal training when every hidden neuron has out-degree one. Additionally, the paper identifies a new class of linear networks—satisfying a novel data throughput condition—that are also efficiently solvable. These results substantially broaden the known tractable regimes for both ReLU and linear-activation networks, improving upon prior findings by Arora et al. and advancing the theoretical foundations of neural network optimization.
📝 Abstract
In spite of the fundamental role of neural networks in contemporary machine learning research, our understanding of the computational complexity of optimally training neural networks remains incomplete even when dealing with the simplest kinds of activation functions. Indeed, while there has been a number of very recent results that establish ever-tighter lower bounds for the problem under linear and ReLU activation functions, less progress has been made towards the identification of novel polynomial-time tractable network architectures. In this article we obtain novel algorithmic upper bounds for training linear- and ReLU-activated neural networks to optimality which push the boundaries of tractability for these problems beyond the previous state of the art. In particular, for ReLU networks we establish the polynomial-time tractability of all architectures where hidden neurons have an out-degree of $1$, improving upon the previous algorithm of Arora, Basu, Mianjy and Mukherjee. On the other hand, for networks with linear activation functions we identify the first non-trivial polynomial-time solvable class of networks by obtaining an algorithm that can optimally train network architectures satisfying a novel data throughput condition.
Problem

Research questions and friction points this paper is trying to address.

neural network training
computational complexity
tractability
ReLU activation
linear activation
Innovation

Methods, ideas, or system contributions that make the work stand out.

computational complexity
polynomial-time tractability
ReLU networks
linear activation
neural network training
🔎 Similar Papers
No similar papers found.
Cornelius Brand
Cornelius Brand
Regensburg University
Algorithms and Complexity
Robert Ganian
Robert Ganian
TU Wien
AlgorithmsParameterized ComplexityArtificial Intelligence
M
Mathis Rocton
Algorithms & Complexity Group, Vienna University of Technology, Favoritenstraße 9-11, 1040 Vienna, Austria