fine-tune vision transformers

Design and implement procedures to adapt pretrained vision transformer architectures (e.g., ViT, DeiT) to new image-based datasets and tasks by choosing which layers or heads to freeze, configuring fine-tuning hyperparameters (learning rates, optimizers, schedulers), modifying input/patch embeddings and classification heads, and applying augmentation and regularization strategies. Measure and analyze model performance and generalization (e.g., accuracy, error rates, cross-domain or zero-shot transfer) to guide iterative fine-tuning and deployment decisions.

fine-tunevisiontransformers

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Diminishing Returns in Self-Supervised Learning

Dec 03, 2025
OB
Oli Bridge
🏛️ University College London

This study investigates the marginal benefits of a three-stage strategy—self-supervised pretraining, intermediate fine-tuning, and downstream task adaptation—for small-scale Vision Transformers (~5M parameters). Motivated by the observation that intermediate fine-tuning may degrade downstream performance due to task misalignment, we propose a systematic ablation framework to assess the impact of varying dataset and objective combinations across stages. Experiments reveal that targeted pretraining substantially improves small-model performance, whereas introducing semantically distant intermediate tasks yields no gain—and often harms performance while wasting compute. The core contribution is the empirical demonstration that, for small ViTs, **the quality of data selection is far more critical than the number of stacked tasks**, challenging conventional assumptions about multi-stage transfer. This finding provides key empirical evidence and methodological guidance for designing efficient, lightweight self-supervised learning paradigms.

Explores diminishing returns in self-supervised learning for small vision transformers.Identifies targeted pre-training and data selection as key for efficient small models.Investigates how intermediate fine-tuning can harm downstream task performance.

CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets

May 13, 2025
AA
Aidar Amangeldi
🏛️ Nazarbayev University

This study investigates the accuracy-efficiency trade-off between CNNs and Vision Transformers (ViTs) on medical (DermaMNIST) and general-purpose (TinyImageNet) image classification tasks. Using ResNet-18 as a baseline, we systematically fine-tune four ViT variants under strict constraints of low latency and parameter count. Methodologically, we introduce a lightweight fine-tuning protocol tailored for resource-constrained deployment. Our key contribution is the first empirical demonstration that minimally adapted ViTs can outperform CNN baselines on cross-domain, small-scale medical data. Results show that ViT-Small achieves a 1.2% accuracy gain on DermaMNIST with 37% fewer parameters and 1.8× faster inference; ViT-Tiny attains 99.4% of ResNet-18’s TinyImageNet accuracy while using only 42% of its parameters. The work establishes ViTs’ viability in low-data medical settings and provides a transferable lightweight adaptation framework for efficient vision model deployment.

Compare CNN and ViT efficiency on medical and general image datasetsEvaluate fine-tuned ViT variants for resource-constrained deploymentReduce inference latency and model complexity with minimal accuracy loss

EA-ViT: Efficient Adaptation for Elastic Vision Transformer

Jul 25, 2025
CZ
Chen Zhu
🏛️ National University of Singapore | Xidian University | University of Toronto | Houmo AI | UCF | Shenzhen Technology University

To address the high energy consumption and inefficiency of repeatedly training Vision Transformers (ViTs) for diverse resource constraints, this paper proposes a single-shot, multi-scale adaptable elastic Vision Transformer framework. The method introduces: (1) a nested elastic architecture that jointly adjusts MLP expansion ratios, number of attention heads, embedding dimension, and network depth; (2) a lightweight learnable router coupled with the NSGA-II multi-objective evolutionary algorithm to identify Pareto-optimal subnetworks; and (3) a progressive curriculum learning strategy to ensure stable knowledge transfer across subnetworks. Evaluated on multiple benchmarks, the framework significantly reduces training overhead while enabling efficient, flexible deployment across heterogeneous hardware platforms. It achieves a favorable trade-off between accuracy and resource efficiency—e.g., FLOPs, latency, and parameter count—without compromising model performance.

Avoiding retraining multiple size-specific ViTs to save resourcesDeploying ViTs for diverse resource constraints efficientlyGenerating adaptable ViT models via elastic architecture and routing

This study addresses the challenge that the effects of module-level features and their interactions on generalization in Vision Transformer (ViT) architectures remain insufficiently isolated and quantified. By analyzing ViT representational structures, we identify feature collapse during initialization and propose quantitative metrics based on feature entropy and minimum eigenvalues. Our analysis reveals the critical roles of token spaces and linear submodules in generalization, systematically validated through multi-scale architectural comparison experiments. Results demonstrate that the proposed surrogate metrics improve correlation rankings with generalization performance by 18%–48%, enabling precise identification of low-compute, high-accuracy ViT architectures. This work provides both theoretical foundations and practical guidance for efficient visual model design.

Architecture DesignFeature InformationGeneralization Behavior

To address the high computational cost, memory footprint, and latency of parameter-efficient fine-tuning (PEFT) for Vision Transformers (ViTs) during inference, this paper proposes a semantic-aware sparse token processing and dense cross-layer Adapter co-architecture. Methodologically: (1) tokens are dynamically sparsified layer-wise based on attention weights and gradient-based semantic importance scores, preserving semantically critical tokens while merging redundant ones; (2) dense shallow-to-deep Adapter modules are introduced to integrate multi-level local features across layers. This work is the first to jointly optimize PEFT efficiency in terms of GFLOPs, GPU memory consumption, and inference latency. Evaluated on VTAB-1K and multiple image/video benchmarks, our approach reduces GFLOPs by 30–38% for ViT-B, significantly lowers inference latency, and achieves state-of-the-art accuracy.

Improves inference efficiency of Vision TransformersMaintains performance despite token sparsification information lossReduces computational and memory overhead in fine-tuning

Latest Papers

What's happening recently
View more

This work challenges the conventional assumption that smoother components are inherently superior by investigating the relationship between component adaptability and smoothness in Vision Transformers under transfer learning. To this end, it introduces “plasticity” as a novel metric quantifying a component’s sensitivity to input perturbations and systematically evaluates the plasticity of attention modules and feed-forward layers through both theoretical analysis and large-scale experiments. The study reveals that components with higher plasticity—i.e., lower smoothness—consistently yield better fine-tuning performance and significantly enhance downstream task accuracy. These findings establish plasticity as a critical lens for assessing adaptation capacity and provide a principled guideline for selecting components in efficient fine-tuning, thereby overturning the long-standing design paradigm centered on smoothness.

finetuningplasticitysmoothness

This work addresses a key limitation in existing parameter-efficient fine-tuning methods for vision Transformers, which adapt each token independently and thereby ignore their intrinsic structural relationships, leading to redundant updates and spatial inconsistency. To overcome this, the authors propose performing adaptation in a hyperedge space: tokens are softly grouped into implicit hyperedges via a soft hypergraph construction, lightweight bottleneck adaptation is applied at the hyperedge level, and updates are propagated back to individual tokens through a hypergraph diffusion mechanism. This approach introduces structured hyperedge adaptation into parameter-efficient fine-tuning for the first time, explicitly modeling inter-token relationships and injecting structural inductive bias while preserving modularity and efficiency. Experiments demonstrate significant performance gains over state-of-the-art methods across multiple vision benchmarks, with particularly notable improvements on tasks requiring structural reasoning, underscoring the critical role of adaptation space design.

hypergraphparameter-efficient fine-tuningstructured relationships

This work addresses the optimization instability and lack of theoretical capacity guidance associated with inserting adapters into frozen vision Transformer backbones during transfer learning. The authors propose the Zero-initialized Residual Low-rank Adapter, which introduces a low-rank bottleneck structure in each Transformer block, with the up-projection layer initialized to zero to ensure that fine-tuning starts identically to the pretrained model, thereby preventing early representation drift. For the first time, the adapter rank is theoretically modeled as a capacity budget tied to the feature shift of downstream tasks, revealing an “elbow”-shaped accuracy gain as rank increases. Experiments across nine datasets and three backbone scales show that the method improves top-1 accuracy by 14.9% on average over training only the classification head, using just 0.92% of the parameters required for full fine-tuning, and outperforms full fine-tuning in 10 out of 15 dataset-backbone combinations.

adapter capacityfrozen-backbone transferlow-rank adapters

This work addresses the challenge of reusing pretrained Softmax attention weights when migrating to linear-complexity attention mechanisms. The authors propose a direct conversion of pretrained Vision Transformers into linear-attention models based on Test-Time Training (TTT) through architectural and representational alignment. They introduce instance normalization and a locality-enhancement module to significantly improve representation consistency. This approach achieves, for the first time, lossless weight transfer from Softmax-based Transformers to linear TTT architectures. With only one hour of fine-tuning on 4×H20 GPUs, the resulting SD3.5-T⁵ model attains 1.32× and 1.47× faster inference at 1K and 2K resolutions, respectively, while preserving generation quality comparable to the original model.

linear attentionrepresentational gapTest-Time Training

ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters

Oct 21, 2025
ZH
Zhiwei Hao
🏛️ Beijing Institute of Technology | City University of Hong Kong | Sun Yat-sen University | Huawei Noah’s Ark Lab

To address the rapid growth in parameters and computational cost during the depth scaling of pretrained Vision Transformers (ViTs), this paper proposes ScaleNet—a training-free, efficient depth-scaling method. ScaleNet extends ViT depth by inserting incremental layers while jointly leveraging inter-layer weight sharing and lightweight learnable adaptation modules (e.g., parallel LoRA-style adapters). Its core innovation lies in mitigating performance degradation induced by depth expansion with negligible parameter overhead (≈0% additional parameters). On ImageNet-1K, a 2× deeper DeiT-Base model achieves a +7.42% top-1 accuracy gain and reduces training time to one-third of conventional retraining. Strong generalization is further validated on downstream object detection tasks. ScaleNet establishes a new paradigm for low-cost, high-performance ViT scaling—enabling substantial depth expansion without architectural redesign or full model retraining.

Efficiently scaling vision transformers with minimal parameter increasesMaintaining performance while expanding model depth through weight sharingReducing computational costs of training large ViT models

Hot Scholars

TK

Tobias Kirschstein

PhD Student @ Technical University of Munich
3D Computer VisionNeural Rendering
SG

Simon Giebenhain

PhD Student, Technical University of Munich
Virtual AvatrsNeural Implicit RepresentationsGeometric Deep Learning3D Scene Understanding
MN

Matthias Nießner

Professor of Computer Science, Technical University of Munich
Computer GraphicsComputer VisionArtificial IntelligenceMachine Learning
GO

Gilberto Ochoa-Ruiz

Tec de Monterrey, CV-inside lab, Advanced AI Research Group
Endoscopic ImagingComputer VisionMedical Image ComputingImage-guided Surgery
AA

Athanasios Angelakis

Department of EpideAmsterdam UMC, Amsterdam Public Health Research Institute, University of
Algebraic Number TheoryDeep LearningMachine Learning