IDInit: A Universal and Stable Initialization Method for Neural Network Training

📅 2025-03-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing initialization methods for residual networks—such as Fixup—struggle to enforce strict identity mappings simultaneously across both the main and shortcut branches, weakening inductive bias and compromising training stability. To address this, we propose IDInit, a fully identity-based initialization scheme. Its core innovation lies in employing padded quasi-identity matrices to overcome rank constraints inherent in non-square weight tensors, thereby enabling dual identity initialization of both the main-path layers and the shortcut branch within each residual block for the first time. We provide theoretical analysis establishing convergence guarantees under SGD, and further enhance robustness via high-order tensor extensions and dynamic dead-neuron compensation. Extensive experiments on large-scale datasets and deep architectures demonstrate that IDInit significantly improves training stability and convergence speed, while achieving superior generalization performance compared to state-of-the-art baselines including Fixup and ReZero.

Technology Category

Computer Vision: Generative Adversarial Networks (GANs) for VisionMachine Learning: Matrix & Tensor MethodsSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Deep neural networks have achieved remarkable accomplishments in practice. The success of these networks hinges on effective initialization methods, which are vital for ensuring stable and rapid convergence during training. Recently, initialization methods that maintain identity transition within layers have shown good efficiency in network training. These techniques (e.g., Fixup) set specific weights to zero to achieve identity control. However, settings of remaining weight (e.g., Fixup uses random values to initialize non-zero weights) will affect the inductive bias that is achieved only by a zero weight, which may be harmful to training. Addressing this concern, we introduce fully identical initialization (IDInit), a novel method that preserves identity in both the main and sub-stem layers of residual networks. IDInit employs a padded identity-like matrix to overcome rank constraints in non-square weight matrices. Furthermore, we show the convergence problem of an identity matrix can be solved by stochastic gradient descent. Additionally, we enhance the universality of IDInit by processing higher-order weights and addressing dead neuron problems. IDInit is a straightforward yet effective initialization method, with improved convergence, stability, and performance across various settings, including large-scale datasets and deep models.
Problem

Research questions and friction points this paper is trying to address.

Improves neural network initialization for stable convergence.
Addresses limitations of identity-preserving initialization methods.
Enhances universality and performance across diverse datasets and models.
Innovation

Methods, ideas, or system contributions that make the work stand out.

IDInit preserves identity in residual networks
Uses padded identity-like matrix for non-square weights
Enhances universality by processing higher-order weights