Score
Analyze and quantify how residual (skip) connections affect signal and gradient propagation and the dynamical stability of deep neural networks by building diagnostics, metrics, and models of the residual stream such as Lyapunov-style stability measures, per-layer jump ratios, and historical-feature signal-to-noise ratios. Use these analyses to relate skip-connection strength to spectral behavior, identify layers that cause gradient explosion/vanishing or quantization instability, and characterize phase-wise activation emergence in the residual pathway.
This work investigates rank collapse in deep Transformers at initialization, where nonlinearities and matrix multiplications degrade representational capacity and training stability. The authors systematically analyze how components within feedforward blocks influence rank preservation across depth, unifying skip connections and normalization mechanisms under a common framework as gradient-based rank-preserving strategies. They reveal a fundamental distinction between Pre-Norm and Post-Norm architectures in terms of rank dynamics and demonstrate that the two-matrix structure and width expansion are critical for maintaining full-rank Jacobians. Through spectral analysis, Jacobian rank tracking, Marchenko–Pastur law modeling, and CIFAR-10 experiments, they establish that the rank of the input–output Jacobian at initialization strongly predicts training success, offering a new principle for deep architecture design grounded in rank evolution.
This study investigates the origins of gradient explosion and vanishing in deep neural networks, with a focus on how residual connections influence gradient propagation dynamics. Drawing upon multiplicative ergodic theory, the work introduces—for the first time—the characterization of Lyapunov exponents due to Furstenberg–Kifer, integrating tools from random matrix theory and modeling the residual architecture to systematically analyze the evolution of gradients. The analysis reveals that residual connections effectively compress the spread of the Lyapunov spectrum, thereby substantially mitigating gradient instability. These findings offer crucial theoretical grounding for the principled design of deep networks, elucidating why residual structures enhance trainability in very deep models.
The scaling factor in residual connections of ResNets critically influences generalization, yet its mechanistic role and robustness across hyperparameter configurations remain poorly understood. Method: We establish the first finite-width field-theoretic framework for ResNets and analytically derive the input response function to characterize signal propagation. Contribution/Results: Our theory reveals that the empirically optimal scaling interval corresponds to the regime of maximal input sensitivity; moreover, the optimal scaling value depends only weakly on network depth and weight variance—explaining its empirical stability across diverse hyperparameter settings. This work provides the first analytical solution for the residual scaling factor and yields interpretable, theoretically grounded guidelines for its selection, thereby bridging empirical practice with rigorous understanding of signal propagation in deep residual networks.
This study investigates the cross-layer dynamical mechanisms of the residual stream (RS) in Transformer models. We model the RS as a continuous dynamical system—introducing dynamical systems theory from neuroscience into large-model interpretability research for the first time. Using orbital stability analysis, dimensionality-reduction visualizations (PCA/t-SNE), and large-scale activation statistics, we identify three key properties: (i) strong cross-layer continuity in the RS; (ii) inter-layer acceleration of evolution, exponential growth in activation density, and unstable periodic orbits; and (iii) curved, attractor-like trajectories in low-dimensional embedding space. Collectively, these findings reveal an underlying dynamical structure in the RS characterized by coexisting stable and unstable regimes. Our work establishes the first dynamical-systems-based theoretical framework for large-model interpretability, grounded in empirical evidence—thereby laying foundational groundwork for an AI neuroscience paradigm.
This work addresses gradient vanishing/exploding issues in deep ResNets as depth increases, systematically investigating training stability mechanisms in ultra-deep networks. Leveraging probabilistic analysis and continuous-limit modeling, we rigorously prove—under standard initialization—that the layer-wise output scaling factor αₗ = 1/√L is the unique nontrivial stable scaling regime; moreover, this scaling induces a Neural Stochastic Differential Equation (Neural SDE) as the continuous limit, challenging the conventional belief that ResNets inherently converge to Neural ODEs. We further uncover a strong coupling between weight regularity and scaling. Supported by theoretical derivation and large-scale initialization/scaling experiments, our framework unifies three distinct regimes: gradient explosion, stable training, and performance degradation. Crucially, we establish that both αₗ and post-training weight smoothness jointly govern generalization performance—both before and after training.
This work addresses the limitation of existing output-confidence–based fault detection methods, which often fail to capture internal errors in neural networks. The authors propose Self-Detecting Neural Networks (SDNN), a novel framework that introduces the concept of “spectral drift” to reveal that erroneous predictions manifest as pronounced multi-scale spectral instabilities in internal activations. Spectral features are extracted via short-time Fourier transform, wavelet decomposition, and statistical moments, and a lightweight detector is trained using curriculum learning to establish an end-to-end learnable internal monitoring mechanism. Evaluated on CIFAR-10, SDNN achieves an AUROC of 79.0 ± 25.3%, outperforming baseline methods such as MaxSoftmax and Energy Score by 25–30 percentage points.
This work addresses the limited understanding of training dynamics in deep neural networks with ReLU activations, particularly regarding how activation patterns evolve during optimization. The study proposes that training unfolds over two distinct time scales: an initial phase characterized by rapid changes in activation patterns, followed by a later phase where weights are fine-tuned within stable activation regions. Leveraging a geometric perspective, the authors develop a theoretical framework for activation pattern stability, supported by measure-theoretic analysis of local stability. They empirically track activation and weight trajectories across fully connected, convolutional, and Transformer architectures, revealing that activation patterns stabilize approximately three times earlier than weight updates converge. This consistent observation—“activations converge first, weights fine-tune later”—provides a foundational insight for staged optimization strategies in deep learning.
This work addresses the well-known difficulty of neural networks in approximating high-frequency functions, where conventional residual connections often fail to effectively capture high-frequency patterns. To overcome this limitation, the authors propose a gradient-enhanced residual connection mechanism that explicitly incorporates input gradients into the skip path for the first time. By forming a learnable convex combination of standard residuals and gradient-based residuals, the method adaptively modulates the network’s reliance on high-frequency information. Theoretically, this design enhances sensitivity to input variations. Empirically, the approach significantly outperforms standard residual networks on high-frequency sinusoidal regression tasks and demonstrates consistent gains in single-image super-resolution, while maintaining competitive performance on standard vision benchmarks such as image classification and segmentation.
This work addresses the long-standing lack of quantitative validation for the analogy between deep neural network forward propagation and renormalization group (RG) flow, particularly the absence of measurable RG order parameters and empirical evidence under controlled inputs. Training pure MLP residual networks on synthetic Markov chains with known spectral properties for masked prediction, the study proposes effective rank as an RG order parameter and combines positional representation tracking with inter-layer kernel drift analysis to quantitatively characterize representational evolution with depth. The findings reveal that effective rank monotonically collapses with depth, but significantly only for inputs with short correlation lengths; layer-wise changes concentrate in a few transition layers, while others converge to fixed-point plateaus. These results demonstrate that MLP residual networks perform input-spectrum-guided selective coarse-graining, offering the first empirical and metric framework substantiating the RG–deep learning analogy.
This work uncovers a negative weight drift phenomenon arising from the coupling between standard loss functions—such as mean squared error and cross-entropy—and positively biased activation functions like ReLU during early training stages, which triggers a sharp increase in activation sparsity and spike-like behavior in intermediate layers. Through theoretical gradient analysis and extensive experiments across diverse architectures—including MLPs, ResNets, Vision Transformers, GPT-nano, and MP-SENet—the study establishes the optimization-theoretic nature and universality of this drift, identifying a critical sparsity threshold near 70% where accuracy precipitously declines. To mitigate these issues, the authors propose squared activation variants with gradient clipping—ReLU² and GELU²—achieving up to 90% activation sparsity in GPT-nano. Notably, clipped ReLU² substantially alleviates spiking, while GELU² yields the lowest validation loss, thereby delineating a clear trade-off boundary between sparsity and model accuracy.