Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficient initialization and suboptimal performance of linear Vision Transformers (ViTs) compared to their Softmax counterparts, identifying the inability to directly reuse attention weights as the core bottleneck. To overcome this, the authors propose a "copy the identical, distill the different" strategy, revealing that MLP weights can be transferred losslessly while attention routing behaviors must be recovered through knowledge distillation. Based on these insights, pretrained Softmax ViT weights are efficiently transferred to linear architectures. Across multiple model scales and datasets, this approach enables linear ViTs to match or even surpass the performance of the original Softmax models. Ultimately, this work establishes a general paradigm for the efficient initialization of linear attention mechanisms.
📝 Abstract
Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and effectively still remains unclear. In this work, we explicitly ask: given that most foundation ViTs are built on the mainstream Softmax attention, can linear ViTs benefit from their pre-trained weights? Recent works on Attention Transfer show that attention is the effective transferable component between Softmax ViTs, suggesting attention alone suffices for such reuse. However, we find the opposite for Softmax-to-linear transfer. The attention weights are operator-specific: copying them barely helps, and is sometimes even worse than random initialization. Instead, the attention's token routing behavior can be recovered through distillation with a proper loss design, letting linear ViTs reduce the gap and even match Softmax ones. In contrast, the MLP weights, which carry the learned representation, are operator-agnostic: they can be transferred by simple direct copying, which already carries most of the benefit of the pre-trained weights. Thus, copying MLPs can serve as an effective foundation for Softmax-to-linear transfer: paired with the distilled attention, linear ViTs eventually close the remaining gap and even surpass Softmax ones. These findings hold consistently across various linear ViT variants, different model sizes, and diverse datasets. We hope this study deepens the understanding of reusing pre-trained weights across attention operators: copy what stays the same and distill what differs, to recover the benefit across the Softmax-to-linear boundary.
Problem

Research questions and friction points this paper is trying to address.

Linear Vision Transformers
Initialization
Pre-trained weights
Attention transfer
Softmax-to-linear
Innovation

Methods, ideas, or system contributions that make the work stand out.

Linear Vision Transformers
Knowledge Distillation
Attention Transfer
Weight Initialization
Token Routing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Huaiyuan Qin
Huaiyuan Qin
Institute for Infocomm Research (I2R), A*STAR, Singapore
Computer VisionDeep Learning
Muli Yang
Muli Yang
Institute for Infocomm Research (I2R), A*STAR, Singapore
Computer VisionMachine LearningOpen-World LearningMultimodal Modeling
G
Gabriel James Goenawan
Institute of Advanced Intelligence and Computing (IAIC), A*STAR, Singapore
Shiqi Huang
Shiqi Huang
PhD Student, Nanyang Technological University
Computer VisionMultimodal LearningRemote Sensing
M
Min Kass Chong
ST Engineering Geo-Insights, Singapore
Wahyu Wiratama
Wahyu Wiratama
ST Engineering Geo-Insights, Singapore
P
Peng Hu
Sichuan University
C
Chen Gong
Shanghai Jiao Tong University
W
Wu Liu
University of Science and Technology of China
X
Xi Peng
Sichuan University
C
Chun Jian Ho
ST Engineering Geo-Insights, Singapore
H
Hongyuan Zhu
Institute of Advanced Intelligence and Computing (IAIC), A*STAR, Singapore