Dissecting Representation Structure in Vision Transformers: A Rigorous Architectural Study

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that the effects of module-level features and their interactions on generalization in Vision Transformer (ViT) architectures remain insufficiently isolated and quantified. By analyzing ViT representational structures, we identify feature collapse during initialization and propose quantitative metrics based on feature entropy and minimum eigenvalues. Our analysis reveals the critical roles of token spaces and linear submodules in generalization, systematically validated through multi-scale architectural comparison experiments. Results demonstrate that the proposed surrogate metrics improve correlation rankings with generalization performance by 18%–48%, enabling precise identification of low-compute, high-accuracy ViT architectures. This work provides both theoretical foundations and practical guidance for efficient visual model design.
📝 Abstract
Representation structure is crucial for understanding Vision Transformer (ViT) architectures and their generalization behavior. However, prior studies neither isolate nor analyze module-level features nor investigate how their interactions contribute to performance estimation. In this work, we conduct the first rigorous analysis of feature information across diverse architectural scales, empirically uncover the relationship between ViT representation and generalization behavior, and leverage these insights to guide efficient ViT design. Our contributions are fivefold: Across diverse architectural scales, 1) We identify feature collapse at initialization, which leads to redundancy, and propose a reduction scheme to mitigate this issue. 2) We quantify feature information using entropy and the minimum eigenvalue, demonstrating that these metrics serve as reliable indicators for generalization prediction. 3) We show that feature in the token space provides a more faithful representation than those in embedding space. 4) We discover an unexpected finding: features produced by linear submodules within ViT layers are critical for the prediction of generalization performance. 5) Our proposed proxy improves the correlation ranking by 18-48% over prior baselines and can effectively identify ViT architectures that achieve higher accuracy at lower or comparable computational cost.
Problem

Research questions and friction points this paper is trying to address.

Vision Transformer
Representation Structure
Generalization Behavior
Feature Information
Architecture Design
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision Transformer
Representation Structure
Feature Collapse
Generalization Prediction
Architecture Design