When less is more: evolving large neural networks from small ones

๐Ÿ“… 2025-01-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF

career value

194K/year
๐Ÿค– AI Summary
Fixed-architecture large neural networks suffer from low computational efficiency and poor generalization. Method: This paper proposes an end-to-end differentiable dynamic scalable feedforward network. Its core innovation is the first formulation of network size as a continuous, differentiable variable explicitly embedded into the loss function, enabling joint gradient-based optimization of both architecture and parameters. A single learnable structural control weight governs node addition and removal during training, allowing the network to grow adaptively from a minimal initial structure to a task-optimal sizeโ€”eliminating the need for post-hoc pruning. The approach integrates scale-aware loss, dynamic computation graph construction, and differentiable structural evolution. Results: Experiments on nonlinear regression and classification tasks demonstrate that the method outperforms static networks of comparable size, automatically converging to smaller, more efficient architectures with superior generalization performance.

Technology Category

Application Category

๐Ÿ“ Abstract
In contrast to conventional artificial neural networks, which are large and structurally static, we study feed-forward neural networks that are small and dynamic, whose nodes can be added (or subtracted) during training. A single neuronal weight in the network controls the network's size, while the weight itself is optimized by the same gradient-descent algorithm that optimizes the network's other weights and biases, but with a size-dependent objective or loss function. We train and evaluate such Nimble Neural Networks on nonlinear regression and classification tasks where they outperform the corresponding static networks. Growing networks to minimal, appropriate, or optimal sizes while training elucidates network dynamics and contrasts with pruning large networks after training but before deployment.
Problem

Research questions and friction points this paper is trying to address.

Adaptive Neural Networks
Size Adjustment
Complex Task Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Network Adjustment
Optimization Algorithm
Real-time Structure Tuning
๐Ÿ”Ž Similar Papers
No similar papers found.
A
Anil Radhakrishnan
Nonlinear Artificial Intelligence Laboratory, Physics Department, North Carolina State University, Raleigh, NC 27607, USA
J
John F. Lindner
Nonlinear Artificial Intelligence Laboratory, Physics Department, North Carolina State University, Raleigh, NC 27607, USA; Physics Department, The College of Wooster, Wooster, OH 44691, USA
S
Scott T. Miller
Nonlinear Artificial Intelligence Laboratory, Physics Department, North Carolina State University, Raleigh, NC 27607, USA
S
Sudeshna Sinha
Indian Institute of Science Education and Research Mohali, Knowledge City, SAS Nagar, Sector 81, Manauli PO 140 306, Punjab, India
W
William L. Ditto
Nonlinear Artificial Intelligence Laboratory, Physics Department, North Carolina State University, Raleigh, NC 27607, USA