design initialization schemes

Design and analyze concrete procedures for setting initial values of model parameters (weights, biases, and random seeds) to break permutation symmetry, control activation and gradient scaling, and satisfy training constraints. This includes developing deterministic or stochastic initialization algorithms, theoretical analyses of variance and signal propagation, and practical strategies to enable stable optimization for specialized architectures such as unrolled or constrained networks.

designinitializationschemes

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

IDInit: A Universal and Stable Initialization Method for Neural Network Training

Mar 06, 2025
YP
Yu Pan
🏛️ Harbin Institute of Technology | The Chinese University of Hong Kong | The Hong Kong Polytechnic University | Meta AI | Fudan University

Existing initialization methods for residual networks—such as Fixup—struggle to enforce strict identity mappings simultaneously across both the main and shortcut branches, weakening inductive bias and compromising training stability. To address this, we propose IDInit, a fully identity-based initialization scheme. Its core innovation lies in employing padded quasi-identity matrices to overcome rank constraints inherent in non-square weight tensors, thereby enabling dual identity initialization of both the main-path layers and the shortcut branch within each residual block for the first time. We provide theoretical analysis establishing convergence guarantees under SGD, and further enhance robustness via high-order tensor extensions and dynamic dead-neuron compensation. Extensive experiments on large-scale datasets and deep architectures demonstrate that IDInit significantly improves training stability and convergence speed, while achieving superior generalization performance compared to state-of-the-art baselines including Fixup and ReZero.

Addresses limitations of identity-preserving initialization methods.Enhances universality and performance across diverse datasets and models.Improves neural network initialization for stable convergence.

Optimized Weight Initialization on the Stiefel Manifold for Deep ReLU Neural Networks

Aug 30, 2025
HL
Hyungu Lee
🏛️ Kyungpook National University

Deep ReLU networks suffer from neuron death (“dying ReLU”) and gradient instability due to suboptimal weight initialization; conventional schemes (e.g., He, Xavier, orthogonal initialization) fail to jointly control pre-activation mean, sparsity, and variance stability—especially in extremely deep architectures. Method: We formulate weight initialization as an optimization problem on the Stiefel manifold, explicitly incorporating the ReLU nonlinearity’s statistical prior, and derive a closed-form family of orthogonal initializations with an efficient sampling strategy. Contribution/Results: Our method theoretically guarantees exact calibration of pre-activation statistics, thereby eliminating neuron death at initialization and mitigating gradient vanishing and variance decay. Empirically, it consistently outperforms state-of-the-art initialization methods across MNIST, Fashion-MNIST, tabular datasets, and few-shot learning tasks. Notably, it maintains training stability and convergence even in networks exceeding 100 layers—where existing approaches collapse.

Controlling activation sparsity through optimized initializationMitigating gradient vanishing in very deep architecturesPreventing neuron inactivation in deep ReLU networks

This work addresses the limited dynamical expressivity of gated recurrent neural networks (RNNs) in reservoir computing, which often arises from suboptimal fixed-weight initialization. By leveraging random matrix theory and phase transition analysis, the authors derive a critical gain criterion applicable to various gated RNN architectures in the infinite-width limit. This criterion guides weight initialization such that the network operates precisely at the edge of chaos—the critical point between ordered and chaotic dynamics. Notably, it establishes the first direct link between the critical weight variance and peak performance in chaotic time series prediction tasks, accurately predicting the optimal initialization gain. The result provides a universal principle for the efficient design of gated RNNs, ensuring maximal computational capacity through principled initialization.

criticalityrandom matrix theoryrecurrent neural networks

Gaussian Pre-Activations in Neural Networks: Myth or Reality?

May 24, 2022
PW
Pierre Wolinski
🏛️ Paris-Dauphine University | PSL University | CNRS | Univ. Grenoble Alpes | Inria | Grenoble INP

This work challenges the conventional infinite-width assumption by investigating whether pre-activation distributions in finite-width neural networks can strictly retain Gaussianity. We derive the first necessary and sufficient constraints ensuring Gaussian pre-activations throughout training, thereby reformulating edge-of-chaos theory with precise analytical characterization. Our framework unifies diverse initialization schemes and systematically evaluates Gaussianity as an optimality criterion for initialization. Leveraging probabilistic propagation modeling, constraints on nonlinear transformations, moment matching, and exact edge-of-chaos analysis, we construct paired families of activation functions and initialization distributions—e.g., tanh with Uniform, SiLU with Gamma—that provably sustain Gaussian pre-activations over extended training in finite-width settings. Theoretical analysis establishes rigorous guarantees, and empirical validation confirms long-term Gaussian fidelity. All code is publicly released to support reproducibility and further extension.

Analyzing constraints for Gaussian pre-activations in neural networksEnsuring Gaussian pre-activations in finite-width neural networksEvaluating desirability of Gaussian pre-activations during initialization

This work investigates the precise influence of initialization on learning dynamics in deep linear networks, specifically how it governs the transition of representation evolution from the “lazy” regime (static representations, constant Neural Tangent Kernel—NTK) to the “rich” regime (dynamic representations, feature learning). Method: We introduce the λ-balanced initialization framework and derive, for the first time, closed-form analytical solutions for the entire training trajectory, jointly characterizing the co-evolution of weights, hidden-layer representations, and the NTK. Contribution/Results: Our theory quantitatively identifies the critical initialization scale separating lazy and rich learning, revealing how initialization determines the paradigm shift. The results yield testable theoretical criteria for continual learning, reverse learning, and transfer learning, thereby addressing a fundamental limitation of existing NTK theory—its inability to model dynamic representation learning.

Derives exact solutions for lambda-balanced initializationsExamines influence of initialization on learning dynamicsExplores impact on learning regimes and Neural Tangent Kernel

Latest Papers

What's happening recently
View more

This work addresses the instability of activation magnitudes in deep Leaky ReLU networks at initialization, particularly in narrow architectures where conventional He and orthogonal initializations fail to ensure stable signal propagation. The authors introduce, for the first time, the Lyapunov exponent as a tool to analyze activation dynamics in deep networks. By leveraging the law of large numbers and the central limit theorem, they characterize the asymptotic behavior of the logarithm of activation norms and explicitly compute the Lyapunov exponent using random matrix theory. Building on this analysis, they propose Lyapunov initialization, which enforces a zero Lyapunov exponent to achieve optimal activation stability. Empirical results demonstrate that this method significantly outperforms existing initialization strategies and substantially improves training performance in narrow, deep networks.

activation stabilitydeep neural networksLeaky ReLU

Function-parameterized neural networks are highly sensitive to initialization, and conventional data-agnostic initialization schemes often fail to capture the structural characteristics of target signals, leading to slow convergence and unstable performance. This work proposes a prior-guided initialization strategy that, for the first time, integrates data-driven spectral priors into both network initialization and architecture design. Specifically, fast Fourier transform (FFT) is employed to extract seasonal priors that inform model depth and initial state, while residual regression is used to parameterize trend components. Without altering the training procedure, the proposed method significantly accelerates convergence, reduces performance variance, and improves computational efficiency across both synthetic and real-world datasets. Notably, it maintains reconstruction accuracy even when using a lower-dimensional encoder, consistently outperforming standard initialization approaches.

convergencedata priorsfunction parameterization

This work addresses the theoretical gap concerning how orthogonal initialization enhances training stability in finite-width neural networks by introducing an analytical framework based on finite-width expansions. By establishing layer-wise recurrence relations for network statistical tensors and generalizing Feynman diagram techniques to arbitrary-order width corrections, the study provides the first complete theoretical explanation of stability under orthogonal initialization. The approach applies to any finite width order and elucidates the mechanism by which tensor statistics saturate and stabilize in the deep-network limit. Theoretical predictions exhibit excellent agreement with Monte Carlo simulations, confirming the framework’s validity and broad applicability.

criticalitydepth stabilityfinite-width networks

This work addresses the barren plateau problem in quantum neural network training caused by poor parameter initialization by proposing a first-moment–based analytical framework. Combining operator concentration theory with numerical experiments, the study systematically evaluates and compares the efficacy of various initialization strategies—including identity, Gaussian, and several shifted or asymmetric distributions. For the first time, it establishes an operator-level criterion for initialization validity, demonstrating that viable initializations avoiding barren plateaus are highly non-unique and form exponentially many inequivalent families. Moreover, the research reveals that initializations with distinct first moments can converge to different local minima, indicating that intelligent initialization effectively transforms the exponential concentration challenge into a selection problem among numerous trainable regions.

barren plateausinitialization strategiesoptimization landscape

Hot Scholars

XG

Xin Geng

School of Computer Science and Engineering, Southeast University
Artificial IntelligencePattern RecognitionMachine Learning
SL

Simon Lucey

Director, Australian Institute for Machine Learning (AIML) + CommBank AI Scholar
Computer VisionMachine LearningArtificial Intelligence
HS

Hemanth Saratchandran

Australian Institute for Machine Learning/Adelaide University + CommBank AI Scholar
MathematicsMachine Learning
LL

Liang Lin

Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis
ST

Supanut Thanasilp

Faculty member, Chulalongkorn University, Thailand
quantum computingquantum machine learningquantum many-body physics