Score
Designs, implements, and analyzes training methods, initialization schemes, activation choices, and architecture modifications that enable stable end-to-end optimization of very deep neural networks without using normalization layers (e.g., LayerNorm). This includes techniques to maintain numerical gradient stability, prevent dead neurons from ReLU-like units, and ensure convergence of deep MLPs using bounded activations.
This work addresses the long-standing lack of theoretical understanding behind the training instability of Transformers, particularly regarding the interplay among normalization schemes, learning rate warmup, and depth scaling. Building from first principles, the authors develop a stability theory grounded in attention sensitivity and the geometric structure of gradient flow. They derive the exact operator norm of the softmax Jacobian for the first time, introduce a sequence-length-independent balance quality factor θ(p) to quantify attention sensitivity, and establish corresponding Lipschitz bounds. Through operator norm analysis, block-wise ∞/RMS geometric modeling, and gradient path characterization, they demonstrate that architectural design—not learned attention patterns—governs stability. The theory successfully explains the efficacy of pre-LayerNorm, DeepNorm’s N⁻¹/⁴ scaling law, and the necessity of learning rate warmup. Empirical validation on a 774M-parameter model confirms θ(p)≈1 throughout training, underscoring that stability is inherently architectural.
Neural networks often suffer from training instability and poor interpretability due to the absence of structural constraints. To address this, we propose a structured hierarchical transformation framework that decomposes each layer into an analytically tractable linear operator and a lightweight residual correction term, explicitly enforcing consistency in information flow while preserving representational capacity. This design significantly improves gradient stability, robustness against adversarial perturbations, and reliability of inter-layer signal propagation—all without modifying the standard backpropagation algorithm, thus ensuring compatibility with diverse architectures and training paradigms. Experiments on synthetic tasks and real-world benchmarks (e.g., ImageNet, WikiText) demonstrate superior gradient condition numbers, reduced input sensitivity, and smoother layer-wise activation responses. The method achieves a principled balance between interpretability and practical performance.
In 16-bit (FP16) training, the Adam optimizer suffers from numerical instability primarily due to its sensitivity to the epsilon hyperparameter—a previously unrecognized root cause of Adam’s failure under low-precision arithmetic. Method: We propose a lightweight, adaptive epsilon dynamic calibration mechanism that requires no additional computation, model modifications, or overhead. It adjusts epsilon in real time based solely on gradient and second-moment statistics, integrating FP16 numerical analysis, gradient scaling, and adaptive stability control while preserving the original Adam framework. Contribution/Results: Our method significantly enhances optimization robustness in FP16 training. Experiments across multiple mainstream models demonstrate convergence speed and final accuracy on par with FP32 training, substantially improved training stability, and zero throughput degradation—achieving high-fidelity low-precision optimization without sacrificing performance or efficiency.
Transformer training instability is often attributed to the lack of theoretical guidance for layer normalization (LN) placement. Method: This paper establishes a unified theoretical framework that quantitatively characterizes how LN positioning affects the growth bound of forward hidden states and the evolution of backward gradient norms, thereby revealing its role in steering optimization toward benign or pathological solutions. Combining theoretical analysis with numerical experiments, we demonstrate that post-residual LN (Post-LN) suppresses state and gradient explosion, while residual scaling coefficients must be co-designed with LN placement. We derive a general stability criterion and use it to optimize scaling strategies. Contribution/Results: Empirical validation shows that principled joint configuration of LN position and scaling significantly improves training stability and final model performance. The framework is extensible to stability analysis of novel Transformer architectures.
This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.
This study investigates the algorithmic stability of over-parameterized deep ReLU homogeneous networks when trained to achieve zero training error via the minimum L²-norm interpolating solution. By integrating algorithmic stability theory, minimum-norm interpolation analysis, and structural properties of neural networks, the work identifies key conditions for stability: robustness to small perturbations in the training data is guaranteed if the network contains a stable subnetwork followed by a low-rank weight layer; conversely, non-low-rank layers may induce instability. These findings offer new theoretical insights into the generalization behavior of minimum-norm interpolating deep networks, shedding light on how architectural choices influence stability and, consequently, generalization in highly over-parameterized regimes.
This work addresses the unclear stability mechanisms of zeroth-order (ZO) optimization methods in deep learning, particularly the lack of theoretical characterization regarding the relationship between step size and the Hessian spectrum. Through mean-square linear stability analysis, we reveal for the first time that the stability condition of ZO methods depends on the full Hessian spectrum rather than solely on its largest eigenvalue—as is typical for first-order methods. We derive a computable stability boundary requiring only the largest eigenvalue and the trace of the Hessian, and further uncover that large step sizes implicitly regularize the Hessian trace in ZO optimization. These theoretical findings apply to ZO-GD, ZO-GDM, and ZO-Adam, and are empirically validated across diverse deep learning tasks, where these methods operate near the predicted stability edge.
This study investigates the role of LayerNorm in pre-normalized recurrent Transformers, focusing on its impact on system stability and memory mechanisms. Through analytical derivations, fixed-point analysis, spectral theory, and from-scratch CPU-level training experiments across six tasks—including ablation studies—the work reveals for the first time that LayerNorm acts as an implicit gain controller within recurrent blocks. This induces a non-normal yet asymptotically contractive Jacobian, establishing that system stability is governed by spectral margin rather than operator norm. The research further clarifies that the carry term primarily stabilizes recurrent dynamics rather than encoding deep memory; genuine memory functionality arises from nonlinear recurrent pathways, with the carry term recruited for memory only under gradient descent in channel-axis-aligned tasks.
This work addresses the challenge of vanishing or exploding activations and gradients in deep neural networks when scale control mechanisms like batch normalization are unavailable—such as in physics-informed neural networks (PINNs)—which often leads to unstable training. The authors propose StableGrad, an optimizer-level inter-layer gradient rescaling mechanism that adaptively corrects weight gradients after backpropagation without altering the forward architecture, adding normalization layers, or employing residual connections. This preserves the physical consistency of both the model output and its derivatives. StableGrad enables stable training without any architectural modifications, significantly improving convergence and solution accuracy in deep PINNs and in ResNet/EfficientNet variants with batch normalization removed, offering a general-purpose, plug-and-play optimization strategy for scenarios where batch normalization is inapplicable.
This work proposes Hierarchical Zeroth-Order Optimization (HZO), a novel approach that overcomes the poor scalability of conventional zeroth-order methods—whose query complexity scales as $O(ML^2)$—to deep neural networks. By introducing a divide-and-conquer strategy along the network depth, HZO departs from the standard layer-wise gradient propagation paradigm and reduces the query complexity to $O(ML \log L)$. The method integrates hierarchical decomposition, rigorous error analysis, and Lipschitz constant control to ensure numerical stability, particularly in the near-unitary regime. Empirical evaluations on CIFAR-10 and ImageNet demonstrate that HZO achieves accuracy comparable to backpropagation, substantially enhancing the scalability and practicality of zeroth-order optimization for deep models.