๐ค AI Summary
This work proposes a novel pruning paradigm grounded in selection dynamics, addressing the limitations of conventional neural network pruning methods that rely on explicit intervention and centralized decision-makingโapproaches ill-suited to the decentralized, stochastic, and path-dependent nature of gradient-based training. By conceptualizing neurons as an evolving population subject to selection pressure, the method defines neuronal fitness through local learning signals, allowing low-fitness components to naturally vanish during training without predefined pruning schedules. Sparsity thus emerges intrinsically as an evolutionary outcome of training rather than being imposed as an external constraint. Evaluated on MNIST with a 768-neuron prunable MLP, the approach achieves 95.5% accuracy at 35% pruning and maintains 88.3โ88.6% accuracy at 50% pruning, closely approaching the 98% performance of the original dense model.
๐ Abstract
Neural networks are commonly trained in highly overparameterized regimes, yet empirical evidence consistently shows that many parameters become redundant during learning. Most existing pruning approaches impose sparsity through explicit intervention, such as importance-based thresholding or regularization penalties, implicitly treating pruning as a centralized decision applied to a trained model. This assumption is misaligned with the decentralized, stochastic, and path-dependent character of gradient-based training. We propose an evolutionary perspective on pruning: parameter groups (neurons, filters, heads) are modeled as populations whose influence evolves continuously under selection pressure. Under this view, pruning corresponds to population extinction: components with persistently low fitness gradually lose influence and can be removed without discrete pruning schedules and without requiring equilibrium computation. We formalize neural pruning as an evolutionary process over population masses, derive selection dynamics governing mass evolution, and connect fitness to local learning signals. We validate the framework on MNIST using a population-scaled MLP (784--512--256--10) with 768 prunable neuron populations. All dynamics reach dense baselines near 98\% test accuracy. We benchmark post-training hard pruning at target sparsity levels (35--50\%): pruning 35\% yields $\approx$95.5\% test accuracy, while pruning 50\% yields $\approx$88.3--88.6\%, depending on the dynamic. These results demonstrate that evolutionary selection produces a measurable accuracy--sparsity tradeoff without explicit pruning schedules during training.