On the Necessity of Attention-FFN Split in Vision Transformers

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses whether the strict architectural separation between attention and feed-forward networks (FFN) in Vision Transformers (ViTs) is necessary, and whether this paradigm constrains the performance ceiling of small-scale models. To investigate this, we break the conventional dichotomy by proposing AttenFeed, a unified module that deeply integrates both functionalities. Building upon this module, we introduce uViT, a novel architecture designed to replace the traditional alternating stacking paradigm, and conduct systematic comparative experiments across multi-scale datasets. Our findings reveal that the strict separation mechanism limits the performance of smaller models due to rigid parameter allocation, thereby validating the theoretical merit of a unified architecture. Ultimately, this work provides new analytical tools and architectural design perspectives for a deeper understanding of the inductive biases inherent in ViTs.
📝 Abstract
The standard Transformer architecture relies on a rigid pattern that alternates Attention and Feed-Forward Network (FFN) layers. Despite its widespread adoption, the inductive bias imposed by this strict separation has not been systematically examined. In this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs). To facilitate this analysis, we introduce the AttenFeed module, a unified component that integrates the functional properties of both Attention and FFN. Based on this module, we devise the unified Vision Transformer (uViT), which replaces the conventional alternating Attention-FFN structure with a sequence of AttenFeed modules. We then use uViT as a control group that relaxes the Attention-FFN dichotomy of the standard ViT and systematically compare the two models across multiple datasets and model scales. Our experiments reveal that the Attention-FFN dichotomy can hinder performance at smaller model scales due to the rigid parameter allocation of ViTs. The AttenFeed module and uViT serve as new analytical tools for understanding the Attention-FFN structure and offer theoretical insights into the heuristically designed architecture of conventional ViTs.
Problem

Research questions and friction points this paper is trying to address.

Vision Transformer
Attention-FFN split
inductive bias
architecture design
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision Transformer
AttenFeed module
Attention-FFN dichotomy
unified Vision Transformer (uViT)
inductive bias
🔎 Similar Papers
No similar papers found.