normalization-free training

Designs, implements, and analyzes training methods, initialization schemes, activation choices, and architecture modifications that enable stable end-to-end optimization of very deep neural networks without using normalization layers (e.g., LayerNorm). This includes techniques to maintain numerical gradient stability, prevent dead neurons from ReLU-like units, and ensure convergence of deep MLPs using bounded activations.

normalization-freetraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the long-standing lack of theoretical understanding behind the training instability of Transformers, particularly regarding the interplay among normalization schemes, learning rate warmup, and depth scaling. Building from first principles, the authors develop a stability theory grounded in attention sensitivity and the geometric structure of gradient flow. They derive the exact operator norm of the softmax Jacobian for the first time, introduce a sequence-length-independent balance quality factor θ(p) to quantify attention sensitivity, and establish corresponding Lipschitz bounds. Through operator norm analysis, block-wise ∞/RMS geometric modeling, and gradient path characterization, they demonstrate that architectural design—not learned attention patterns—governs stability. The theory successfully explains the efficacy of pre-LayerNorm, DeepNorm’s N⁻¹/⁴ scaling law, and the necessity of learning rate warmup. Empirical validation on a 774M-parameter model confirms θ(p)≈1 throughout training, underscoring that stability is inherently architectural.

attention sensitivitygradient flowsoftmax Jacobian

Neural networks often suffer from training instability and poor interpretability due to the absence of structural constraints. To address this, we propose a structured hierarchical transformation framework that decomposes each layer into an analytically tractable linear operator and a lightweight residual correction term, explicitly enforcing consistency in information flow while preserving representational capacity. This design significantly improves gradient stability, robustness against adversarial perturbations, and reliability of inter-layer signal propagation—all without modifying the standard backpropagation algorithm, thus ensuring compatibility with diverse architectures and training paradigms. Experiments on synthetic tasks and real-world benchmarks (e.g., ImageNet, WikiText) demonstrate superior gradient condition numbers, reduced input sensitivity, and smoother layer-wise activation responses. The method achieves a principled balance between interpretability and practical performance.

Lack of stable learning in neural networksNeed for interpretable neural network behaviorUnconstrained affine transformations causing training instability

Stable Adam Optimization for 16-bit Neural Networks Training

Jul 30, 2023
JY
Juyoung Yun
🏛️ Stony Brook University | MODULABS

In 16-bit (FP16) training, the Adam optimizer suffers from numerical instability primarily due to its sensitivity to the epsilon hyperparameter—a previously unrecognized root cause of Adam’s failure under low-precision arithmetic. Method: We propose a lightweight, adaptive epsilon dynamic calibration mechanism that requires no additional computation, model modifications, or overhead. It adjusts epsilon in real time based solely on gradient and second-moment statistics, integrating FP16 numerical analysis, gradient scaling, and adaptive stability control while preserving the original Adam framework. Contribution/Results: Our method significantly enhances optimization robustness in FP16 training. Experiments across multiple mainstream models demonstrate convergence speed and final accuracy on par with FP32 training, substantially improved training stability, and zero throughput degradation—achieving high-fidelity low-precision optimization without sacrificing performance or efficiency.

Address numerical instability in 16-bit neural trainingIdentify epsilon hyperparameter as instability source in AdamPropose modified Adam optimizer for stable 16-bit training

Stability of Transformers under Layer Normalization

Oct 10, 2025
KK
Kelvin Kan
🏛️ UCLA | UT Austin | UNC Chapel Hill | SRI International | UMass Amherst

Transformer training instability is often attributed to the lack of theoretical guidance for layer normalization (LN) placement. Method: This paper establishes a unified theoretical framework that quantitatively characterizes how LN positioning affects the growth bound of forward hidden states and the evolution of backward gradient norms, thereby revealing its role in steering optimization toward benign or pathological solutions. Combining theoretical analysis with numerical experiments, we demonstrate that post-residual LN (Post-LN) suppresses state and gradient explosion, while residual scaling coefficients must be co-designed with LN placement. We derive a general stability criterion and use it to optimize scaling strategies. Contribution/Results: Empirical validation shows that principled joint configuration of LN position and scaling significantly improves training stability and final model performance. The framework is extensible to stability analysis of novel Transformer architectures.

Analyzing forward and backward stability in Transformers with layer normalizationProviding theoretical framework to evaluate architectural modifications' stabilityStudying how normalization placement affects training dynamics and gradients

This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.

deep learning librarieseducational gapfundamental understanding

Latest Papers

What's happening recently
View more

This study investigates the algorithmic stability of over-parameterized deep ReLU homogeneous networks when trained to achieve zero training error via the minimum L²-norm interpolating solution. By integrating algorithmic stability theory, minimum-norm interpolation analysis, and structural properties of neural networks, the work identifies key conditions for stability: robustness to small perturbations in the training data is guaranteed if the network contains a stable subnetwork followed by a low-rank weight layer; conversely, non-low-rank layers may induce instability. These findings offer new theoretical insights into the generalization behavior of minimum-norm interpolating deep networks, shedding light on how architectural choices influence stability and, consequently, generalization in highly over-parameterized regimes.

algorithmic stabilitydeep ReLU networksgeneralization error

This work addresses the unclear stability mechanisms of zeroth-order (ZO) optimization methods in deep learning, particularly the lack of theoretical characterization regarding the relationship between step size and the Hessian spectrum. Through mean-square linear stability analysis, we reveal for the first time that the stability condition of ZO methods depends on the full Hessian spectrum rather than solely on its largest eigenvalue—as is typical for first-order methods. We derive a computable stability boundary requiring only the largest eigenvalue and the trace of the Hessian, and further uncover that large step sizes implicitly regularize the Hessian trace in ZO optimization. These theoretical findings apply to ZO-GD, ZO-GDM, and ZO-Adam, and are empirically validated across diverse deep learning tasks, where these methods operate near the predicted stability edge.

Deep learningHessian spectrumImplicit regularization

This study investigates the role of LayerNorm in pre-normalized recurrent Transformers, focusing on its impact on system stability and memory mechanisms. Through analytical derivations, fixed-point analysis, spectral theory, and from-scratch CPU-level training experiments across six tasks—including ablation studies—the work reveals for the first time that LayerNorm acts as an implicit gain controller within recurrent blocks. This induces a non-normal yet asymptotically contractive Jacobian, establishing that system stability is governed by spectral margin rather than operator norm. The research further clarifies that the carry term primarily stabilizes recurrent dynamics rather than encoding deep memory; genuine memory functionality arises from nonlinear recurrent pathways, with the carry term recruited for memory only under gradient descent in channel-axis-aligned tasks.

gain controlLayerNormLooped Transformers

This work addresses the challenge of vanishing or exploding activations and gradients in deep neural networks when scale control mechanisms like batch normalization are unavailable—such as in physics-informed neural networks (PINNs)—which often leads to unstable training. The authors propose StableGrad, an optimizer-level inter-layer gradient rescaling mechanism that adaptively corrects weight gradients after backpropagation without altering the forward architecture, adding normalization layers, or employing residual connections. This preserves the physical consistency of both the model output and its derivatives. StableGrad enables stable training without any architectural modifications, significantly improving convergence and solution accuracy in deep PINNs and in ResNet/EfficientNet variants with batch normalization removed, offering a general-purpose, plug-and-play optimization strategy for scenarios where batch normalization is inapplicable.

Batch Normalizationdeep neural networksgradient scale control

This work proposes Hierarchical Zeroth-Order Optimization (HZO), a novel approach that overcomes the poor scalability of conventional zeroth-order methods—whose query complexity scales as $O(ML^2)$—to deep neural networks. By introducing a divide-and-conquer strategy along the network depth, HZO departs from the standard layer-wise gradient propagation paradigm and reduces the query complexity to $O(ML \log L)$. The method integrates hierarchical decomposition, rigorous error analysis, and Lipschitz constant control to ensure numerical stability, particularly in the near-unitary regime. Empirical evaluations on CIFAR-10 and ImageNet demonstrate that HZO achieves accuracy comparable to backpropagation, substantially enhancing the scalability and practicality of zeroth-order optimization for deep models.

Computational complexityDeep neural networksNon-differentiable objectives

Hot Scholars

FM

Fandong Meng

WeChat AI, Tencent
Machine TranslationNatural Language Processing
CS

Chenze Shao

Tencent
Machine TranslationNatural Language ProcessingDeep Learning
OS

Ortal Senouf

École Polytechnique Fédérale de Lausanne (EPFL)
IC

Irit Chelly

PhD student in Computer Science, Ben-Gurion University
machine learningdeep learningprobabilistic graphical modelsspatial transformations
AB

Adam Block

Columbia University
Machine Learning Theory