residual connection design

Designs, builds, and analyzes skip/residual pathways and gating mechanisms that preserve or selectively bypass feature identity across layers, connect modules, and integrate outputs between components. Works include choosing where to place residual links, how to gate or scale them, and how they affect gradient/feature flow and training stability for compacted or modularized networks.

residualconnectiondesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work investigates rank collapse in deep Transformers at initialization, where nonlinearities and matrix multiplications degrade representational capacity and training stability. The authors systematically analyze how components within feedforward blocks influence rank preservation across depth, unifying skip connections and normalization mechanisms under a common framework as gradient-based rank-preserving strategies. They reveal a fundamental distinction between Pre-Norm and Post-Norm architectures in terms of rank dynamics and demonstrate that the two-matrix structure and width expansion are critical for maintaining full-rank Jacobians. Through spectral analysis, Jacobian rank tracking, Marchenko–Pastur law modeling, and CIFAR-10 experiments, they establish that the rank of the input–output Jacobian at initialization strongly predicts training success, offering a new principle for deep architecture design grounded in rank evolution.

depth scalinggradient rankrank collapse

This study investigates whether explicitly exposing routing mechanisms in Transformers is sufficient to achieve mechanistic interpretability. To this end, the authors propose Block Attention Residuals, which represent cross-layer information routing as observable tensors during forward propagation and enable causal intervention to analyze their functional roles. Experiments based on the Qwen3 architecture demonstrate that meaningful local routing patterns emerge only when the routing structure is actively involved in training optimization. Crucially, the magnitude of routing weights does not directly reflect causal importance, necessitating intervention-based validation of interpretability hypotheses. The work identifies three characteristic local routing patterns and reveals that segments with the largest routing weights do not necessarily contribute the most causally, establishing that explicit exposure of routing mechanisms is necessary but insufficient for mechanistic interpretability.

attention residualscausal probinginterpretability

ResNets Are Deeper Than You Think

Jun 17, 2025
CH
Christian H.X. Ali Mehmeti-Gopel
🏛️ Johannes-Gutenberg University

This work investigates whether residual connections merely constitute a reparameterization of feedforward networks or instead confer fundamentally distinct functional representational capacity. To isolate architectural effects from optimization confounds, we conduct controlled post-training analysis—comparing generalization performance between residual and equivalent-depth feedforward networks under identical weight initialization, fixed parameters, and shared training dynamics. We find that residual architectures consistently outperform their feedforward counterparts, demonstrating that their superiority stems from intrinsic differences in function space rather than mere optimization convenience. Based on this, we propose a novel “variable-depth” inductive bias: residual structures implicitly enable cross-depth information reuse, better aligning with the hierarchical structure of natural data. This study provides the first causally controlled empirical evidence that residual networks operate within a distinct function space, thereby revealing the structural origin of their generalization advantage.

Residual connections enable different function space than feedforwardResNets outperform feedforward networks in generalizationResNets' inductive bias aligns better with natural data

Three Mechanisms of Feature Learning in a Linear Network

Jan 13, 2024
YX
Yizhou Xu
🏛️ Abdus Salam International Center for Theoretical Physics | Massachusetts Institute of Technology | NTT Research

This work investigates how neural network width governs training dynamics. For single-hidden-layer linear networks, we derive the first exact analytical solution of learning dynamics at arbitrary finite width, unifying the characterization of the two-phase evolution—kernel learning and feature learning—and establishing a complete phase diagram parameterized by width, layer-wise learning rates, and initialization scale. Methodologically, we integrate analytical dynamical systems analysis, phase-diagram modeling, and empirical validation on nonlinear networks. Crucially, we identify three novel mechanisms operative during the feature-learning phase: alignment learning, de-alignment learning, and rescaling learning—each transcending the conventional kernel-method paradigm. These theoretical insights are empirically reproduced in realistic deep networks, offering a new conceptual framework for understanding training dynamics and designing adaptive optimization algorithms. (138 words)

Analyzes learning dynamics in neural networksExplores hyperparameter impact on training trajectoriesIdentifies feature learning mechanisms in networks

Field theory for optimal signal propagation in ResNets

May 12, 2023
KF
Kirsten Fischer
🏛️ Jülich Research Centre | RWTH Aachen University

The scaling factor in residual connections of ResNets critically influences generalization, yet its mechanistic role and robustness across hyperparameter configurations remain poorly understood. Method: We establish the first finite-width field-theoretic framework for ResNets and analytically derive the input response function to characterize signal propagation. Contribution/Results: Our theory reveals that the empirically optimal scaling interval corresponds to the regime of maximal input sensitivity; moreover, the optimal scaling value depends only weakly on network depth and weight variance—explaining its empirical stability across diverse hyperparameter settings. This work provides the first analytical solution for the residual scaling factor and yields interpretable, theoretically grounded guidelines for its selection, thereby bridging empirical practice with rigorous understanding of signal propagation in deep residual networks.

Analyzing signal propagation in residual networks with scaling parametersDeriving optimal scaling values for improved generalization performanceExplaining universality of scaling parameters across network hyperparameters

Latest Papers

What's happening recently
View more

This work addresses the insufficient robustness of current models under natural image corruptions, particularly their vulnerability in safety-critical scenarios. The authors present the first explicit characterization of internal robust computational pathways within neural networks, revealing a consistent attenuation of robust features across layers. To counteract this degradation, they propose a novel “Suppress and Diversify” mechanism that is architecture-agnostic, parameter-free, and incurs zero overhead at test time. This approach dynamically selects and diversifies symmetry-preserving robust pathways to enhance overall model robustness. Extensive experiments across eight benchmarks demonstrate that the method consistently improves performance across diverse vision tasks, backbone architectures, and complex real-world conditions, highlighting its strong generalizability and scalability.

corruption robustnessinternal robustnessmodel robustness

This work addresses the limitation of conventional residual connections, which sum sublayer updates with fixed coefficients and cannot dynamically assess the reliability of proposed updates. To overcome this, the authors propose Review Residuals—a novel mechanism that explicitly incorporates conditional dependence on the proposed update within the residual gating function. By employing a learnable sigmoid gate conditioned on two inputs via RMSNorm, the method dynamically scales the residual term while preserving the identity additive structure, thereby balancing training stability and representational capacity. The approach integrates seamlessly into standard Transformers and demonstrates statistically significant improvements (p<0.05) over both standard residual connections and Highway gating in models of 590M parameters and larger, with performance gains increasing with model scale. It also enables stable training of extremely deep networks.

gated residualsreliabilityresidual connections

Hot Scholars

BZ

Bo Zheng

Researcher, Alibaba Group
AINetworkE-Commerce
HJ

Hyunwoo J. Kim

Korea Advanced Institute of Science & Technology (KAIST)
Deep LearningMachine LearningComputer VisionGraph Neural Networks
JX

Jian Xu

Senior Director, Ad Platform, Alibaba Group
Computational AdvertisingMachine LearningData MiningData Privacy
CL

Changchun Li

Jilin University
Text ClassificationTopic ModelingWeakly Supervised LearningPartial Label Learning