gated skip connections

Design, implement, or analyze neural network skip-connection modules that insert gated pathways between layers to selectively route feature maps and control information flow; gates may be local, global, or a combination to allow selective forwarding of reliable details and to decouple functions such as noise suppression from texture preservation. Work includes choices of skip-connection topology and parameterization, including variants that use Dirac (identity) initialization for direct identity routing and other skip-connection design decisions.

gatedskipconnections

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates whether a single-hidden-layer multilayer perceptron (MLP) of fixed width can, through parameter reconfiguration, implicitly absorb its skip connections and thereby eliminate the explicit residual structure. By combining homogeneity analysis, linearization arguments, matrix algebra, and perspectives from function spaces and probability measures, the work establishes necessary and sufficient conditions under which such absorption is possible. It proves that for activation functions like ReLU and GELU, absorption holds only under specific algebraic constraints, whereas for ReLU², SwiGLU, and GeGLU, it is generally infeasible under generic weight configurations. These findings imply that, at equal width, MLPs with and without skip connections typically represent function classes that are almost everywhere disjoint, and single-block absorption is achievable only under non-generic conditions.

activation functionfunction classMLP

Computing Linear Regions in Neural Networks with Skip Connections

Sep 18, 2025
JJ
Johnny Joyce
🏛️ University of Illinois at Chicago

This work investigates the mechanistic impact of skip connections on the number of linear regions in deep neural networks. For piecewise-linear activation networks, we propose an exact modeling framework grounded in tropical algebra, wherein forward propagation is represented as compositions of tropical polynomials, and design an efficient algorithm to analytically compute the total number of linear regions across the entire network. Theoretical analysis and empirical evaluation demonstrate that skip connections substantially increase the number of linear regions, and this growth exhibits structural interpretability: it stems from mitigating gradient degradation and enhancing functional expressivity, thereby improving training stability and suppressing overfitting. To our knowledge, this is the first work to establish a quantitative link between skip connections and the geometric complexity of linear regions—measured via their count—providing novel evidence from tropical geometry for the generalization advantage of residual architectures.

Applying tropical geometry to analyze piecewise linear activation functionsComputing linear regions in neural networks with skip connectionsInvestigating overfitting problems and benefits of skip connections

This work investigates rank collapse in deep Transformers at initialization, where nonlinearities and matrix multiplications degrade representational capacity and training stability. The authors systematically analyze how components within feedforward blocks influence rank preservation across depth, unifying skip connections and normalization mechanisms under a common framework as gradient-based rank-preserving strategies. They reveal a fundamental distinction between Pre-Norm and Post-Norm architectures in terms of rank dynamics and demonstrate that the two-matrix structure and width expansion are critical for maintaining full-rank Jacobians. Through spectral analysis, Jacobian rank tracking, Marchenko–Pastur law modeling, and CIFAR-10 experiments, they establish that the rank of the input–output Jacobian at initialization strongly predicts training success, offering a new principle for deep architecture design grounded in rank evolution.

depth scalinggradient rankrank collapse

Enhancing Feature Fusion of U-like Networks with Dynamic Skip Connections

Sep 18, 2025
YC
Yue Cao
🏛️ Sichuan University | University of Maryland

U-Net’s static skip connections suffer from two key limitations: (i) inter-feature constraints—lacking content-aware cross-layer feature interaction—and (ii) intra-feature constraints—insufficient multi-scale feature aggregation. To address these, we propose the Dynamic Skip Connection Module (DSCM), which jointly incorporates test-time training (TTT) and dynamic multi-scale kernels (DMSK) to enable content-adaptive fusion of high- and low-level features and global-context-guided multi-scale interaction. DSCM is architecture-agnostic and seamlessly integrates into diverse U-shaped backbones—including CNNs, Transformers, and Mamba-based models. Evaluated across multiple medical image segmentation benchmarks, DSCM consistently improves segmentation accuracy while demonstrating strong generalizability and robustness to domain shifts and annotation noise.

Addressing insufficient multi-scale feature interaction modelingEnhancing feature fusion in U-like medical image networksOvercoming static inter-feature constraints in skip connections

CHNNet: An Artificial Neural Network With Connected Hidden Neurons

May 17, 2023
RS
Rafiad Sadat Shahir
🏛️ BRAC University

Conventional feedforward neural networks are constrained by strictly layered architectures, where neurons within a layer are disconnected, thereby limiting lateral interaction and intra-layer information integration. Method: This paper proposes CHNNet, the first fully connected artificial neural network architecture systematically incorporating intra-hidden-layer lateral connections. It employs intra-layer weight sharing and optimized gradient propagation paths to enhance dynamic inter-neuron interaction and intra-layer integration. Contribution/Results: We provide a theoretical proof that CHNNet’s convergence rate is strictly superior to that of standard feedforward networks. Empirical evaluations across multiple benchmark tasks demonstrate that CHNNet significantly accelerates convergence while improving generalization stability and training efficiency. The core innovation lies in breaking the hierarchical constraint to establish a provably convergent intra-hidden-layer connectivity paradigm—marking a fundamental departure from traditional architectural assumptions in deep learning.

Achieving faster convergence compared to conventional feedforward networksEnabling lateral interactions for enhanced intra-layer information integrationIntroducing intra-layer neuron connections to overcome hierarchical limitations

Latest Papers

What's happening recently
View more

This work addresses a critical gap in existing approaches to asymmetric path planning, where a disconnect exists between representation and decision-making: while encoding captures pairwise costs, decoding relies solely on node-context compatibility and neglects directed edge-level transition information. To bridge this gap, the authors propose an edge-aware decoding mechanism that explicitly incorporates three candidate-specific signals—namely, the current directed edge cost, the return-to-depot closing cost, and lightweight static lookahead information—without modifying the backbone architecture. Integrated with SVD/Sinkhorn-based asymmetric backbones, the method enables zero-shot generalization from ATSP-100 to larger instances, reducing the optimality gap on ATSP-1000 from 4.13% to 2.73% and consistently improving performance on the Asymmetric Capacitated Vehicle Routing Problem (ACVRP).

ATSPdirected transitionedge-aware decoding

This work proposes a learnable inter-filter connectivity mechanism that replaces the fixed pointwise nonlinear activations in conventional convolutional neural networks with a parameterized, universal connection function embedded within convolutional layers. By enabling adaptive interactions among filters, the approach overcomes the limitations of traditional fixed logical operations—such as multiplication or minimum selection—and allows the network to automatically optimize its connectivity strategy through end-to-end training. Experimental results demonstrate that this method significantly improves classification accuracy, confirming its effectiveness in enhancing both model expressivity and generalization capability.

convolutional neural networksfilter connectionslearnable connections

Traditional neural networks struggle to model higher-order topological structures such as nodes, edges, faces, and hyperedges, often losing critical information when simplifying data into graphs or sequences. This work proposes a general U-Net architecture grounded in combinatorial complexes, replacing conventional spatial scales with “rank” as the hierarchical dimension. Cross-scale feature propagation is achieved through cells, incidence maps, and rank-wise pathways, while a bottleneck support ratio is introduced to quantify compression severity. The framework enables cohomological lifting and rank-matched skip connections across diverse topological domains, revealing the structural role of skip connections under high compression. Empirical results demonstrate that the model achieves state-of-the-art average accuracy on six out of eight node classification datasets and four out of five hypergraph benchmarks, with particularly pronounced gains on heterophilic graphs.

combinatorial complexesencoder-decoder architecturehigher-order structure

This work addresses the challenges of optimization collapse and topological constraints that hinder performance gains in deep traditional logic gate networks. To overcome these limitations, the authors propose Input-Anchored Logic Gate Networks (IALGN), which establish stable information pathways by directly connecting each layer’s logic gates to the original input. The approach integrates input-anchored topology, skip-biased initialization, straight-through gradient estimation, and Random-k anchor relaxation to enforce a strict path-depth hierarchy, ensuring effective utilization of input information even in very deep architectures. Using this framework, the authors successfully train logic gate networks exceeding 100 layers on MNIST, CIFAR-10, and CIFAR-100, achieving simultaneous improvements in depth and accuracy and significantly outperforming existing architectures.

computational pathsdepth scalabilityinformation access

This work addresses the inefficiency of multiply-accumulate (MAC) operations that bottleneck neural network inference on resource-constrained edge-device CPUs. To overcome this limitation, the authors propose a novel approach that equivalently transforms a neural network into a decision tree, extracts decision paths leading to constant leaf nodes, and compresses them into a control-flow-dominated logic structure composed primarily of if-else statements. This transformation effectively bypasses the majority of MAC computations and represents the first efficient conversion from dataflow-centric neural networks to control-flow-centric logical programs—aligning naturally with CPU execution characteristics. Evaluated on a RISC-V CPU simulator, the method achieves up to a 14.9% reduction in inference latency while preserving model accuracy exactly.

CPU efficiencyedge computinglogic flows

Hot Scholars

SW

Saad Wazir

Korea Advanced Institute of Science & Technology (KAIST)
Artificial IntelligenceComputer VisionCloud ComputingWeb Development
RM

Racheal Mukisa

Kent State University
Artificial IntelligenceMachine LearningHealth Informatics
ZP

Zhenming Peng

Professor,University of Electronic Science and Technology of China
Image ProcessingMachine LearningObject DetectionRemote Sensing
LQ

Lin Qi

Ocean University of China
Computer VisionAI for Oceanography