Score
Designs, implements, and evaluates methods for stabilizing training in architectures that route signals through residual detail paths by controlling their initial sign and amplitude; this includes initializing detail sources with a negative bias and applying detached RMS matching to detail signals so their root-mean-square is matched without creating gradient feedback. These techniques prevent destabilizing training dynamics and enable stable optimization at larger network depths.
This work addresses the lack of theoretical stability guarantees in deep residual networks, whose current designs rely heavily on empirical tuning. We propose a sublinear growth principle that ensures stable training without normalization layers by constraining the input-amplitude exponent \( q \leq 1 \) of the residual block’s velocity field, establishing the first certifiable architectural framework with provable stability. This condition is shown to be both necessary and sufficient, directly linking stability to the input-amplitude exponent of architectural primitives for the first time. Leveraging ordinary differential equations and Hamilton–Jacobi–Bellman optimal control theory, we construct a function space for velocity fields and develop an exponent calculus encompassing five fundamental operations. Experiments demonstrate that reducing the supercritical Mamba block’s exponent from \( q = 5 \) to \( q = 1 \) yields consistently stable and efficient training on both Mamba and PatchTST architectures.
This work addresses the instability commonly encountered in large-scale Transformer training, where abrupt loss spikes and divergence lead to wasted computational resources. The authors propose an efficient online curvature estimation algorithm that combines power iteration with warm starts and Hessian-vector products to dynamically track the dominant eigenvalue of the preconditioned Hessian in real time. Furthermore, they introduce a novel architecture warm-up mechanism based on progressively increasing network depth, which actively controls curvature guided by the Edge of Stability theory. Experiments on billion-parameter Transformers demonstrate that the proposed approach significantly enhances training stability, effectively mitigates divergence, and maintains the original convergence rate without degradation.
This work addresses the limited understanding of training dynamics in deep neural networks with ReLU activations, particularly regarding how activation patterns evolve during optimization. The study proposes that training unfolds over two distinct time scales: an initial phase characterized by rapid changes in activation patterns, followed by a later phase where weights are fine-tuned within stable activation regions. Leveraging a geometric perspective, the authors develop a theoretical framework for activation pattern stability, supported by measure-theoretic analysis of local stability. They empirically track activation and weight trajectories across fully connected, convolutional, and Transformer architectures, revealing that activation patterns stabilize approximately three times earlier than weight updates converge. This consistent observation—“activations converge first, weights fine-tune later”—provides a foundational insight for staged optimization strategies in deep learning.
This work addresses the feedback loops that arise after model deployment due to performativity—wherein the model’s predictions influence the data distribution—particularly under strong interventions where the convergence behavior of retraining remains poorly understood. The paper introduces the “stable signal principle,” positing that the prediction target contains an intrinsic component independent of the model (e.g., inherent item quality), and leverages this insight to analyze the dynamics of regularized repeated risk minimization. Theoretically, it establishes that as long as a non-zero stable signal exists, retraining converges geometrically to its direction, even when model-induced effects dominate. This reveals a novel role for regularization in mitigating performative feedback and extends the framework to nonlinear, heterogeneous, and time-varying settings—including language models—thereby explaining the observed stability of training on generated data.
This work investigates the dynamic evolution of the largest Hessian eigenvalue during deep neural network training, focusing on the fundamental distinction between “sharpening–stabilization” behaviors under full-batch versus mini-batch optimization. Addressing high-dimensional regimes, we integrate random matrix theory, Neural Tangent Kernel (NTK) analysis, and Hessian spectral modeling to propose the novel “conservative sharpening” theory, which characterizes a mini-batch–induced deceleration mechanism in curvature growth. We establish, for the first time, that the stability boundary of mini-batch optimization is governed by the trace of the NTK—not by conventional Hessian eigenvalues—and empirically validate both the existence and controllability of this stochastic stability edge. Our findings yield new theoretical principles for understanding optimization dynamics in overparameterized settings and provide a rigorous foundation for designing robust training algorithms.
This work addresses the lack of principled exploitation of rescaling symmetries in ReLU neural networks, which often leads to unstable training dynamics. Building upon the path-lifting framework, the authors propose a geometrically motivated rescaling criterion that aligns the kernel in path space with a reference kernel by minimizing this criterion. This approach systematically integrates rescaling symmetry into the optimization process to enhance training efficiency. Notably, it introduces the first conditioning-based rescaling strategy grounded in the geometric structure of path space. Numerical experiments demonstrate that the proposed method significantly accelerates the training of ReLU networks, highlighting the benefits of leveraging geometric insights for symmetry-aware optimization.
This study investigates how weight decay enhances training stability in deep learning through a unified framework combining dynamical systems analysis, the Edge of Stability (EoS) theory, the Neural Tangent Kernel (NTK) perspective, and mathematical modeling. The authors demonstrate that weight decay induces architecture-dependent phase transitions in both CNNs and MLPs, rooted in the global alignment between parameter vectors and curvature gradients. This mechanism effectively suppresses asymptotic sharpness and modulates oscillations in optimization trajectories. Furthermore, the work reveals that conventional curvature-based thresholds derived under convexity assumptions fail under regularization, thereby establishing weight decay as a nontrivial yet essential regulator of stable training dynamics.
This work addresses the instability of existing heuristic post-training guidance methods for generative flows by formulating guidance as a Lyapunov control problem. It establishes, for the first time, an equivalence between guided flow matching and Lyapunov control, and introduces a pseudo-projection operator that unifies diverse guidance strategies—such as classifier-, reward-, or energy-based guidance—under both model-driven and data-driven settings. This framework provides explicit stability guarantees while preserving computational efficiency. Empirical results demonstrate that the proposed approach significantly improves sample quality, guidance fidelity, and robustness across a range of tasks, including synthetic data generation, image inverse problems, reinforcement learning planning, and energy-based modeling.
本文提出HybridAL方法,通过监控稳定信号自适应地在主动学习中从重新训练切换到微调,以节省时间并保持性能。
This study addresses the unclear dynamics and optimizer roles in deep learning training prior to reaching the Edge of Stability (EoS). Through dense learning rate sweeps, Hessian eigenvalue tracking, and gradient alignment analysis, it proposes a DC-gain normalization method to unify training trajectories across different optimizers. The work reveals the differential effects of key factors under low, medium, and high learning rates. It demonstrates that DC normalization eliminates optimizer discrepancies at low-to-medium learning rates, confirming that progressive sharpening originates from the model rather than the optimizer. Furthermore, it establishes that at high learning rates, the optimizer determines the timing of EoS entry. Overall, this research systematically clarifies the distinct roles of optimizers, learning rates, and models in shaping training dynamics.