Score
Design, implement, or analyze neural network skip-connection modules that insert gated pathways between layers to selectively route feature maps and control information flow; gates may be local, global, or a combination to allow selective forwarding of reliable details and to decouple functions such as noise suppression from texture preservation. Work includes choices of skip-connection topology and parameterization, including variants that use Dirac (identity) initialization for direct identity routing and other skip-connection design decisions.
This study investigates whether a single-hidden-layer multilayer perceptron (MLP) of fixed width can, through parameter reconfiguration, implicitly absorb its skip connections and thereby eliminate the explicit residual structure. By combining homogeneity analysis, linearization arguments, matrix algebra, and perspectives from function spaces and probability measures, the work establishes necessary and sufficient conditions under which such absorption is possible. It proves that for activation functions like ReLU and GELU, absorption holds only under specific algebraic constraints, whereas for ReLU², SwiGLU, and GeGLU, it is generally infeasible under generic weight configurations. These findings imply that, at equal width, MLPs with and without skip connections typically represent function classes that are almost everywhere disjoint, and single-block absorption is achievable only under non-generic conditions.
This work investigates the mechanistic impact of skip connections on the number of linear regions in deep neural networks. For piecewise-linear activation networks, we propose an exact modeling framework grounded in tropical algebra, wherein forward propagation is represented as compositions of tropical polynomials, and design an efficient algorithm to analytically compute the total number of linear regions across the entire network. Theoretical analysis and empirical evaluation demonstrate that skip connections substantially increase the number of linear regions, and this growth exhibits structural interpretability: it stems from mitigating gradient degradation and enhancing functional expressivity, thereby improving training stability and suppressing overfitting. To our knowledge, this is the first work to establish a quantitative link between skip connections and the geometric complexity of linear regions—measured via their count—providing novel evidence from tropical geometry for the generalization advantage of residual architectures.
This work investigates rank collapse in deep Transformers at initialization, where nonlinearities and matrix multiplications degrade representational capacity and training stability. The authors systematically analyze how components within feedforward blocks influence rank preservation across depth, unifying skip connections and normalization mechanisms under a common framework as gradient-based rank-preserving strategies. They reveal a fundamental distinction between Pre-Norm and Post-Norm architectures in terms of rank dynamics and demonstrate that the two-matrix structure and width expansion are critical for maintaining full-rank Jacobians. Through spectral analysis, Jacobian rank tracking, Marchenko–Pastur law modeling, and CIFAR-10 experiments, they establish that the rank of the input–output Jacobian at initialization strongly predicts training success, offering a new principle for deep architecture design grounded in rank evolution.
U-Net’s static skip connections suffer from two key limitations: (i) inter-feature constraints—lacking content-aware cross-layer feature interaction—and (ii) intra-feature constraints—insufficient multi-scale feature aggregation. To address these, we propose the Dynamic Skip Connection Module (DSCM), which jointly incorporates test-time training (TTT) and dynamic multi-scale kernels (DMSK) to enable content-adaptive fusion of high- and low-level features and global-context-guided multi-scale interaction. DSCM is architecture-agnostic and seamlessly integrates into diverse U-shaped backbones—including CNNs, Transformers, and Mamba-based models. Evaluated across multiple medical image segmentation benchmarks, DSCM consistently improves segmentation accuracy while demonstrating strong generalizability and robustness to domain shifts and annotation noise.
Conventional feedforward neural networks are constrained by strictly layered architectures, where neurons within a layer are disconnected, thereby limiting lateral interaction and intra-layer information integration. Method: This paper proposes CHNNet, the first fully connected artificial neural network architecture systematically incorporating intra-hidden-layer lateral connections. It employs intra-layer weight sharing and optimized gradient propagation paths to enhance dynamic inter-neuron interaction and intra-layer integration. Contribution/Results: We provide a theoretical proof that CHNNet’s convergence rate is strictly superior to that of standard feedforward networks. Empirical evaluations across multiple benchmark tasks demonstrate that CHNNet significantly accelerates convergence while improving generalization stability and training efficiency. The core innovation lies in breaking the hierarchical constraint to establish a provably convergent intra-hidden-layer connectivity paradigm—marking a fundamental departure from traditional architectural assumptions in deep learning.
This work addresses a critical gap in existing approaches to asymmetric path planning, where a disconnect exists between representation and decision-making: while encoding captures pairwise costs, decoding relies solely on node-context compatibility and neglects directed edge-level transition information. To bridge this gap, the authors propose an edge-aware decoding mechanism that explicitly incorporates three candidate-specific signals—namely, the current directed edge cost, the return-to-depot closing cost, and lightweight static lookahead information—without modifying the backbone architecture. Integrated with SVD/Sinkhorn-based asymmetric backbones, the method enables zero-shot generalization from ATSP-100 to larger instances, reducing the optimality gap on ATSP-1000 from 4.13% to 2.73% and consistently improving performance on the Asymmetric Capacitated Vehicle Routing Problem (ACVRP).
This work proposes a learnable inter-filter connectivity mechanism that replaces the fixed pointwise nonlinear activations in conventional convolutional neural networks with a parameterized, universal connection function embedded within convolutional layers. By enabling adaptive interactions among filters, the approach overcomes the limitations of traditional fixed logical operations—such as multiplication or minimum selection—and allows the network to automatically optimize its connectivity strategy through end-to-end training. Experimental results demonstrate that this method significantly improves classification accuracy, confirming its effectiveness in enhancing both model expressivity and generalization capability.
Traditional neural networks struggle to model higher-order topological structures such as nodes, edges, faces, and hyperedges, often losing critical information when simplifying data into graphs or sequences. This work proposes a general U-Net architecture grounded in combinatorial complexes, replacing conventional spatial scales with “rank” as the hierarchical dimension. Cross-scale feature propagation is achieved through cells, incidence maps, and rank-wise pathways, while a bottleneck support ratio is introduced to quantify compression severity. The framework enables cohomological lifting and rank-matched skip connections across diverse topological domains, revealing the structural role of skip connections under high compression. Empirical results demonstrate that the model achieves state-of-the-art average accuracy on six out of eight node classification datasets and four out of five hypergraph benchmarks, with particularly pronounced gains on heterophilic graphs.
This work addresses the challenges of optimization collapse and topological constraints that hinder performance gains in deep traditional logic gate networks. To overcome these limitations, the authors propose Input-Anchored Logic Gate Networks (IALGN), which establish stable information pathways by directly connecting each layer’s logic gates to the original input. The approach integrates input-anchored topology, skip-biased initialization, straight-through gradient estimation, and Random-k anchor relaxation to enforce a strict path-depth hierarchy, ensuring effective utilization of input information even in very deep architectures. Using this framework, the authors successfully train logic gate networks exceeding 100 layers on MNIST, CIFAR-10, and CIFAR-100, achieving simultaneous improvements in depth and accuracy and significantly outperforming existing architectures.
This work addresses the inefficiency of multiply-accumulate (MAC) operations that bottleneck neural network inference on resource-constrained edge-device CPUs. To overcome this limitation, the authors propose a novel approach that equivalently transforms a neural network into a decision tree, extracts decision paths leading to constant leaf nodes, and compresses them into a control-flow-dominated logic structure composed primarily of if-else statements. This transformation effectively bypasses the majority of MAC computations and represents the first efficient conversion from dataflow-centric neural networks to control-flow-centric logical programs—aligning naturally with CPU execution characteristics. Evaluated on a RISC-V CPU simulator, the method achieves up to a 14.9% reduction in inference latency while preserving model accuracy exactly.