ResNets Are Deeper Than You Think

📅 2025-06-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates whether residual connections merely constitute a reparameterization of feedforward networks or instead confer fundamentally distinct functional representational capacity. To isolate architectural effects from optimization confounds, we conduct controlled post-training analysis—comparing generalization performance between residual and equivalent-depth feedforward networks under identical weight initialization, fixed parameters, and shared training dynamics. We find that residual architectures consistently outperform their feedforward counterparts, demonstrating that their superiority stems from intrinsic differences in function space rather than mere optimization convenience. Based on this, we propose a novel “variable-depth” inductive bias: residual structures implicitly enable cross-depth information reuse, better aligning with the hierarchical structure of natural data. This study provides the first causally controlled empirical evidence that residual networks operate within a distinct function space, thereby revealing the structural origin of their generalization advantage.

Technology Category

Machine Learning: Deep Neural Architectures and Foundation ModelsComputer Vision: Representation Learning for VisionNatural Language Processing: Learning & Optimization for NLP

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web dataSocial Networks and Social Media: Social media analysis through the lenses of networks
📝 Abstract
Residual connections remain ubiquitous in modern neural network architectures nearly a decade after their introduction. Their widespread adoption is often credited to their dramatically improved trainability: residual networks train faster, more stably, and achieve higher accuracy than their feedforward counterparts. While numerous techniques, ranging from improved initialization to advanced learning rate schedules, have been proposed to close the performance gap between residual and feedforward networks, this gap has persisted. In this work, we propose an alternative explanation: residual networks do not merely reparameterize feedforward networks, but instead inhabit a different function space. We design a controlled post-training comparison to isolate generalization performance from trainability; we find that variable-depth architectures, similar to ResNets, consistently outperform fixed-depth networks, even when optimization is unlikely to make a difference. These results suggest that residual connections confer performance advantages beyond optimization, pointing instead to a deeper inductive bias aligned with the structure of natural data.
Problem

Research questions and friction points this paper is trying to address.

ResNets outperform feedforward networks in generalization
Residual connections enable different function space than feedforward
ResNets' inductive bias aligns better with natural data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Residual connections enhance trainability and accuracy
Variable-depth architectures outperform fixed-depth networks
ResNets have deeper inductive bias for natural data
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Christian H.X. Ali Mehmeti-Gopel
Department of Computer Science, Johannes-Gutenberg University, Mainz, Germany
M
Michael Wand
Department of Computer Science, Johannes-Gutenberg University, Mainz, Germany