Score
Designs, builds, and analyzes skip/residual pathways and gating mechanisms that preserve or selectively bypass feature identity across layers, connect modules, and integrate outputs between components. Works include choosing where to place residual links, how to gate or scale them, and how they affect gradient/feature flow and training stability for compacted or modularized networks.
This work investigates rank collapse in deep Transformers at initialization, where nonlinearities and matrix multiplications degrade representational capacity and training stability. The authors systematically analyze how components within feedforward blocks influence rank preservation across depth, unifying skip connections and normalization mechanisms under a common framework as gradient-based rank-preserving strategies. They reveal a fundamental distinction between Pre-Norm and Post-Norm architectures in terms of rank dynamics and demonstrate that the two-matrix structure and width expansion are critical for maintaining full-rank Jacobians. Through spectral analysis, Jacobian rank tracking, Marchenko–Pastur law modeling, and CIFAR-10 experiments, they establish that the rank of the input–output Jacobian at initialization strongly predicts training success, offering a new principle for deep architecture design grounded in rank evolution.
This study investigates whether explicitly exposing routing mechanisms in Transformers is sufficient to achieve mechanistic interpretability. To this end, the authors propose Block Attention Residuals, which represent cross-layer information routing as observable tensors during forward propagation and enable causal intervention to analyze their functional roles. Experiments based on the Qwen3 architecture demonstrate that meaningful local routing patterns emerge only when the routing structure is actively involved in training optimization. Crucially, the magnitude of routing weights does not directly reflect causal importance, necessitating intervention-based validation of interpretability hypotheses. The work identifies three characteristic local routing patterns and reveals that segments with the largest routing weights do not necessarily contribute the most causally, establishing that explicit exposure of routing mechanisms is necessary but insufficient for mechanistic interpretability.
This work investigates whether residual connections merely constitute a reparameterization of feedforward networks or instead confer fundamentally distinct functional representational capacity. To isolate architectural effects from optimization confounds, we conduct controlled post-training analysis—comparing generalization performance between residual and equivalent-depth feedforward networks under identical weight initialization, fixed parameters, and shared training dynamics. We find that residual architectures consistently outperform their feedforward counterparts, demonstrating that their superiority stems from intrinsic differences in function space rather than mere optimization convenience. Based on this, we propose a novel “variable-depth” inductive bias: residual structures implicitly enable cross-depth information reuse, better aligning with the hierarchical structure of natural data. This study provides the first causally controlled empirical evidence that residual networks operate within a distinct function space, thereby revealing the structural origin of their generalization advantage.
This work investigates how neural network width governs training dynamics. For single-hidden-layer linear networks, we derive the first exact analytical solution of learning dynamics at arbitrary finite width, unifying the characterization of the two-phase evolution—kernel learning and feature learning—and establishing a complete phase diagram parameterized by width, layer-wise learning rates, and initialization scale. Methodologically, we integrate analytical dynamical systems analysis, phase-diagram modeling, and empirical validation on nonlinear networks. Crucially, we identify three novel mechanisms operative during the feature-learning phase: alignment learning, de-alignment learning, and rescaling learning—each transcending the conventional kernel-method paradigm. These theoretical insights are empirically reproduced in realistic deep networks, offering a new conceptual framework for understanding training dynamics and designing adaptive optimization algorithms. (138 words)
The scaling factor in residual connections of ResNets critically influences generalization, yet its mechanistic role and robustness across hyperparameter configurations remain poorly understood. Method: We establish the first finite-width field-theoretic framework for ResNets and analytically derive the input response function to characterize signal propagation. Contribution/Results: Our theory reveals that the empirically optimal scaling interval corresponds to the regime of maximal input sensitivity; moreover, the optimal scaling value depends only weakly on network depth and weight variance—explaining its empirical stability across diverse hyperparameter settings. This work provides the first analytical solution for the residual scaling factor and yields interpretable, theoretically grounded guidelines for its selection, thereby bridging empirical practice with rigorous understanding of signal propagation in deep residual networks.
This work addresses the insufficient robustness of current models under natural image corruptions, particularly their vulnerability in safety-critical scenarios. The authors present the first explicit characterization of internal robust computational pathways within neural networks, revealing a consistent attenuation of robust features across layers. To counteract this degradation, they propose a novel “Suppress and Diversify” mechanism that is architecture-agnostic, parameter-free, and incurs zero overhead at test time. This approach dynamically selects and diversifies symmetry-preserving robust pathways to enhance overall model robustness. Extensive experiments across eight benchmarks demonstrate that the method consistently improves performance across diverse vision tasks, backbone architectures, and complex real-world conditions, highlighting its strong generalizability and scalability.
This work addresses the limitation of conventional residual connections, which sum sublayer updates with fixed coefficients and cannot dynamically assess the reliability of proposed updates. To overcome this, the authors propose Review Residuals—a novel mechanism that explicitly incorporates conditional dependence on the proposed update within the residual gating function. By employing a learnable sigmoid gate conditioned on two inputs via RMSNorm, the method dynamically scales the residual term while preserving the identity additive structure, thereby balancing training stability and representational capacity. The approach integrates seamlessly into standard Transformers and demonstrates statistically significant improvements (p<0.05) over both standard residual connections and Highway gating in models of 590M parameters and larger, with performance gains increasing with model scale. It also enables stable training of extremely deep networks.