fine-tune transformers

Design, build, and evaluate workflows that adapt pretrained transformer architectures to specific downstream tasks by preparing task-labeled datasets, selecting or adding task-specific output heads and loss functions, and running fine-tuning (full or parameter-efficient) with hyperparameter search and regularization; measure and optimize models for target evaluation metrics (for example macro-F1) and for deployment concerns such as per-query inference cost.

fine-tunetransformers

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

PEFT A2Z: Parameter-Efficient Fine-Tuning Survey for Large Language and Vision Models

Apr 19, 2025
NJ
Nusrat Jahan Prottasha
🏛️ University of Central Florida | Khulna University of Engineering & Technology | Delineate Inc. | Universiti Tenaga Nasional | University of South Florida | Daffodil International University | Anymate Me

Full-parameter fine-tuning of large language models (LLMs) and vision-language models (VLMs) suffers from prohibitive computational costs, overfitting, and catastrophic forgetting. Method: We propose the first structured taxonomy of parameter-efficient fine-tuning (PEFT), encompassing additive, selective, reparameterized, hybrid, and unified frameworks, and conduct the first standardized cross-modal (language/vision) and cross-task (understanding/generation) evaluation. Contribution/Results: Through theoretical analysis and multi-domain transfer experiments, we comprehensively benchmark mainstream PEFT methods—including LoRA, Adapter, and Prompt Tuning—demonstrating up to 95% GPU memory reduction and substantial computational savings while retaining over 90% of full fine-tuning performance. We further uncover fundamental trade-offs among robustness, scalability, and interpretability across PEFT paradigms, establishing a methodological foundation and empirical basis for efficient, reliable, and generalizable multimodal model adaptation.

Addresses high computational costs of full fine-tuning for large modelsAnalyzes challenges like overfitting and inefficiency in traditional approachesExplores parameter-efficient methods to adapt models with minimal updates

Must-Read Papers

Most classic and influential ideas
View more

Diminishing Returns in Self-Supervised Learning

Dec 03, 2025
OB
Oli Bridge
🏛️ University College London

This study investigates the marginal benefits of a three-stage strategy—self-supervised pretraining, intermediate fine-tuning, and downstream task adaptation—for small-scale Vision Transformers (~5M parameters). Motivated by the observation that intermediate fine-tuning may degrade downstream performance due to task misalignment, we propose a systematic ablation framework to assess the impact of varying dataset and objective combinations across stages. Experiments reveal that targeted pretraining substantially improves small-model performance, whereas introducing semantically distant intermediate tasks yields no gain—and often harms performance while wasting compute. The core contribution is the empirical demonstration that, for small ViTs, **the quality of data selection is far more critical than the number of stacked tasks**, challenging conventional assumptions about multi-stage transfer. This finding provides key empirical evidence and methodological guidance for designing efficient, lightweight self-supervised learning paradigms.

Explores diminishing returns in self-supervised learning for small vision transformers.Identifies targeted pre-training and data selection as key for efficient small models.Investigates how intermediate fine-tuning can harm downstream task performance.

Update Your Transformer to the Latest Release: Re-Basin of Task Vectors

May 28, 2025
FR
Filippo Rinaldi
🏛️ AImageLab | University of Modena and Reggio Emilia | Sapienza University of Rome

When pretrained foundation models are updated, existing fine-tuned models become obsolete, necessitating efficient knowledge transfer mechanisms—especially under constraints of no access to original training data or computational resources for retraining. Method: We propose a training-free, data-free fine-tuning knowledge transfer method, the first to adapt the *re-basin* paradigm to Transformer architectures. Our approach introduces a spectral-theory-driven, two-level weight rearrangement scheme: (i) attention head permutation, (ii) intra-head parameter alignment, and (iii) task-vector rebasing. Crucially, it resolves structural inconsistencies induced by residual connections and multi-head attention. Results: The method enables zero-shot, zero-step adaptation of legacy fine-tuned models to updated pretrained backbones across vision and language tasks, fully recovering original performance without any gradient updates—eliminating the need for costly retraining.

Apply model re-basin principles to Transformer architecturesEnable data-free knowledge transfer between pretrained backbonesTransfer fine-tuning to new model releases without retraining

This work addresses the challenge of explicitly controlling overfitting during fine-tuning of pretrained Transformers by formulating it as a bilevel optimization regularized framework. It introduces, for the first time, a linear programming–driven local search mechanism that leverages validation gradients and training Hessian information from a warm-up phase to construct a validation-aware descent direction. This enables joint, task-adaptive optimization of both model parameters and regularization hyperparameters without requiring repeated full retraining. Experimental results demonstrate significant reductions in test perplexity on GPT-2 Small and WikiText-2, with particularly pronounced gains in settings prone to overfitting. The approach consistently yields stable improvements across diverse layer configurations and regularization settings.

bilevel optimizationlocal searchoverfitting

This paper addresses the lack of a unified, reproducible framework for adaptive inference in Transformer models. To this end, we introduce AdaptBench—the first end-to-end open-source benchmark for adaptive inference—integrating three core techniques: progressive token pruning, sparse attention, and dynamic early exiting, enabling input-adaptive computation. The framework fully automates the GLUE evaluation pipeline, including data preprocessing, low-overhead timing, CSV-based logging, ablation studies, and joint accuracy–latency assessment. All components are modular, well-documented, and script-ready, significantly enhancing reproducibility and cross-method comparability. On SST-2, AdaptBench achieves marginally higher accuracy than optimized DistilBERT at substantially lower latency, demonstrating the efficacy and practicality of dynamic computation in low-latency NLP applications. By providing a standardized, extensible evaluation infrastructure, AdaptBench establishes a foundational benchmark for future research on adaptive Transformers.

Evaluates dynamic computation trade-offs between accuracy and latencyProvides reproducible framework for adaptive transformer research benchmarkingUnifies three adaptive efficiency techniques for transformer inference optimization

PETAH: Parameter Efficient Task Adaptation for Hybrid Transformers in a resource-limited Context

Oct 23, 2024
MA
Maximilian Augustin
🏛️ Tübingen AI Center | University of Tübingen | Reality Labs Research | FAIR at Meta

Existing hybrid CNN-Transformer vision models lack efficient multi-task adaptation methods under resource-constrained settings. Method: This paper proposes PETAH, a parameter-efficient task-adaptation framework that— for the first time—integrates parameter-efficient fine-tuning (e.g., LoRA variants) into hybrid backbones and couples it with channel- and layer-aware structured pruning to jointly optimize storage and computation. Lightweight adapter modules are co-designed with hybrid backbone fine-tuning. Results: On multiple vision tasks—including classification—PETAH significantly outperforms mainstream ViT adaptation approaches: it reduces trainable parameters by 37%, accelerates mobile inference by 1.8×, and maintains or improves accuracy by up to 0.5%. This work establishes a new paradigm for deploying lightweight, multi-task vision models in edge and resource-limited environments.

Efficient task adaptation for hybrid transformersParameter reduction for resource-constrained applicationsStorage-friendly multi-tasking models for mobile hardware

Latest Papers

What's happening recently
View more

This work addresses the challenge of sparse and noisy observational data in few-shot, large-scale decision-making problems by introducing the pretraining–fine-tuning paradigm to this setting for the first time. The authors propose a problem-specific Transformer architecture that leverages domain knowledge to generate synthetic data for pretraining, followed by fine-tuning on a small amount of real-world data. Theoretically, they establish the first non-asymptotic generalization error bound, elucidating the synergistic mechanism between pretraining and fine-tuning and revealing a scaling law for fine-tuning. Empirically, high-capacity models effectively learn structural priors from synthetic data and adapt efficiently to real environments, with decision performance improving significantly as the instance scale grows.

cross-instance learninglarge-scale optimizationnoisy observations

This work addresses the high computational cost of conventional fine-tuning for instance segmentation, which typically requires updating a large fraction (40–55%) of parameters in large pre-trained models. To improve parameter efficiency, the study explores parameter-efficient fine-tuning (PEFT) methods, introducing LoRA into deformable attention mechanisms for the first time and systematically evaluating the trade-offs between performance and efficiency based on the number and placement of adapters within the Transformer architecture. Experimental results demonstrate that by fine-tuning only 1–6% of the model parameters, the proposed approach matches or even surpasses the performance of full fine-tuning across four benchmark datasets. These findings validate the efficacy and feasibility of PEFT for instance segmentation and further reveal that its effectiveness is influenced by dataset complexity and model architecture.

instance segmentationlarge pretrained modelsparameter-efficient fine-tuning

Large language models (LLMs) face significant challenges in task adaptation under resource-constrained and closed-source API settings, where conventional parameter-efficient fine-tuning (PEFT) methods are inapplicable due to their reliance on direct model parameter access and high computational overhead. Method: This paper proposes a lightweight, parameter-free knowledge injection framework that enables task-specific adaptation without accessing the LLM’s internal parameters. Its core innovation is the “Specialized Small Model (SSM) Collaboration Paradigm,” integrating knowledge distillation from the LLM, distribution-aware task modeling, and zero-parameter coupling between the SSM and the LLM. Contribution/Results: Experiments demonstrate that our approach matches PEFT-level performance across diverse downstream tasks while reducing GPU memory consumption by over 90% and inference latency by 85%. Crucially, it operates entirely within black-box API environments—requiring no model weights, gradients, or architectural access—thus enabling seamless integration with proprietary, closed-source LLM APIs.

Adapts large models to tasks without accessing their parameters.Enhances performance on specific distributions using small models.Reduces resource costs for fine-tuning in constrained environments.

Existing theoretical frameworks struggle to explain why larger-scale pre-trained models substantially reduce sample complexity on downstream tasks. This work proposes a novel theoretical framework—termed “caulking”—inspired by parameter-efficient fine-tuning methods such as adapters, low-rank adaptation, and partial fine-tuning. It establishes, for the first time, a provable relationship between the scale of pre-trained models and the sample complexity of downstream tasks. By rigorously linking stronger pre-training capabilities to reduced data requirements in transfer learning, this study not only addresses a critical gap in current theoretical understanding but also provides a solid foundation for empirically observed scaling laws, demonstrating that enhanced pre-training capacity can significantly decrease the number of samples needed for effective downstream adaptation.

downstream taskspre-trained modelssample complexity

Parameter-Efficient Multi-Task Learning via Progressive Task-Specific Adaptation

Sep 23, 2025
NG
Neeraj Gangwar
🏛️ University of Illinois Urbana-Champaign | Amazon

To address generalization degradation caused by task interference and negative transfer in multi-task learning, this paper proposes Progressive Task-Specific Adaptation (PTSA). PTSA hierarchically introduces lightweight adapter modules atop a shared backbone network and incorporates a gradient-similarity-based dynamic task clustering mechanism to adaptively allocate shared versus task-specific parameters, enabling parameter-efficient fine-tuning. Crucially, we embed gradient similarity measurement directly into the Swin Transformer architecture. On the PASCAL-Context and NYUD-v2 multi-task benchmarks, PTSA achieves superior performance using only 20% of the trainable parameters required by full fine-tuning—outperforming both the full fine-tuning baseline and existing state-of-the-art methods. The approach simultaneously enhances model efficiency, improves task decoupling, and strengthens cross-task generalization capability.

Addressing task interference and negative transfer in multi-task learningDeveloping parameter-efficient adaptation for multiple downstream tasksOptimizing task-specific learning while minimizing trainable parameters

Hot Scholars

GS

Grigori Sidorov

Professor of Computational Linguistics, Instituto Politécnico Nacional (IPN), Mexico
Computational LinguisticsNatural Language ProcessingArtificial IntelligenceMachine Learning
AH

Ali Hamdi

Computer Science, MSA University
Computer VisionDeep LearningText Mining
YA

Yehudit Aperstein

Afeka Academic College of Engineering
Robust AI for Intelligent SystemsGenerative AIMulti-agent Systems
RJ

Raviraj Joshi

Indian Institute of Technology Madras
computer sciencemachine learningnatural language processing
OK

Olga Kolesnikova

Centro de Investigación en Computación (CIC) del Instituto Politécnico Nacional, Mexico
Artificial IntelligenceNatural Language ProcessingLinguistics