How Controlling the Variance can Improve Training Stability of Sparsely Activated DNNs and CNNs

📅 2026-02-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of training instability and limited representational capacity commonly encountered in sparsely activated deep neural networks under high sparsity. By analyzing the Gaussian process characteristics of intermediate layers, the study reveals the critical role of initial weight variance in sparse network performance and proposes a novel mechanism that increases this initial variance. The approach integrates Edge-of-Chaos initialization, a shifted and truncated ReLU activation function (CReLU_{τ,m}), and Gaussian process theory to substantially enhance both training stability and expressive power. Empirical results demonstrate that the method achieves up to 90% sparsity in hidden layers of both DNNs and CNNs while maintaining near-full-precision accuracy, offering a promising pathway toward efficient, low-energy neural network training.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionCognitive Modeling & Cognitive Systems: Neural Spike CodingNatural Language Processing: Learning & Optimization for NLP

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web dataResponsible Web: Measurement, analysis, and circumvention of Web censorship
📝 Abstract
The intermediate layers of deep networks can be characterised as a Gaussian process, in particular the Edge-of-Chaos (EoC) initialisation strategy prescribes the limiting covariance matrix of the Gaussian process. Here we show that the under-utilised chosen variance of the Gaussian process is important in the training of deep networks with sparsity inducing activation, such as a shifted and clipped ReLU, $\text{CReLU}_{\tau,m}(x)=\min(\max(x-\tau,0),m)$. Specifically, initialisations leading to larger fixed Gaussian process variances, allow for improved expressivity with activation sparsity as large as 90% in DNNs and CNNs, and generally improve the stability of the training process. Enabling full, or near full, accuracy at such high levels of sparsity in the hidden layers suggests a promising mechanism to reduce the energy consumption of machine learning models involving fully connected layers.
Problem

Research questions and friction points this paper is trying to address.

sparsely activated DNNs
training stability
Gaussian process variance
activation sparsity
Edge-of-Chaos
Innovation

Methods, ideas, or system contributions that make the work stand out.

variance control
sparsely activated DNNs
Edge-of-Chaos initialization
CReLU
training stability
🔎 Similar Papers
No similar papers found.