fine-tune llms

Design, implement, and evaluate methods to adapt pretrained large or small language model weights and outputs for specific tasks or behaviors, including full-weight fine-tuning, instruction-tuning, adding and training task-specific heads (e.g., regression or classification), transfer learning, and reinforcement-learning-based fine-tuning. Build and apply parameter-efficient and quantization-aware approaches (LoRA, quantized LoRA/QLoRA, low-rank adapters), and perform multilingual or cross-lingual tuning on translated or language-balanced datasets while measuring task metrics, representational drift, adversarial robustness/compliance, and compute/memory trade-offs.

fine-tunellms

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-3.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$198K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Dynamic Adaptation of LoRA Fine-Tuning for Efficient and Task-Specific Optimization of Large Language Models

Jan 24, 2025
CW
Chihang Wang
🏛️ New York University | University of Minnesota | Tulane University | The Chinese University of Hong Kong | Stevens Institute of Technology

To address the inefficiency of fine-tuning large language models (LLMs) under resource constraints and their poor adaptability to multimodal tasks, this paper proposes Dynamic LoRA—a lightweight, input-aware fine-tuning method. Its core innovations are a layer-wise importance reallocation mechanism and a differentiable adapter architecture based on low-rank decomposition. By leveraging gradient-sensitive scoring and input-feature-distribution-guided parameter reweighting, Dynamic LoRA unifies task-specific customization with cross-task generalization. Compared to static LoRA, it achieves 88.1% accuracy and 87.3% F1 score on the GLUE benchmark while incurring only a 0.1% increase in computational overhead—significantly improving the efficiency–performance trade-off. This work establishes a novel paradigm for resource-sensitive, multimodal LLM adaptation.

Efficient Large Language ModelsMultimodal Task FlexibilityResource Optimization

Accurate and Efficient Fine-Tuning of Quantized Large Language Models Through Optimal Balance

Jul 24, 2024
AS
Ao Shen
🏛️ National University of Defense Technology

Quantized LLM fine-tuning suffers from “balance mismatch”: LoRA adapters exhibit high input/output complexity yet low effective trainability, leading to underfitting; subsequent low-precision conversion further degrades performance. Method: We propose Q-BaRA and QA-HiRA—Q-BaRA enhances adapter effectiveness via structural simplification and joint rank optimization; QA-HiRA introduces a single-matrix high-rank adapter coupled with block-wise quantization alignment, enabling lossless integration of fine-tuned parameters into 4/8-bit inference models. Contribution/Results: This work is the first to identify and characterize this imbalance, establishing a new quantization-aware end-to-end fine-tuning paradigm. Evaluated on LLaMA and LLaMA2, our methods significantly outperform state-of-the-art approaches in accuracy, while preserving parameter count and computational cost during training and incurring zero additional deployment overhead.

Address performance degradation in quantized LLM fine-tuningBalance adapter complexity and trainability to reduce underfittingEnable efficient low-precision deployment with quantization-aware fine-tuning

LoRA achieves efficiency in single-task fine-tuning but struggles with cross-task knowledge transfer and relies heavily on task-specific data in multi-task settings. To address this, we propose MeTA-LoRA, a two-stage low-rank adaptation framework: (1) a task-specific stage, where lightweight adapters are learned from minimal per-task samples; and (2) a meta-adaptation stage, where a shared adapter is updated via gradient aggregation across tasks, enabling implicit knowledge transfer. To our knowledge, this is the first work to systematically integrate multi-task knowledge transfer into the LoRA paradigm. Experiments across multi-task and multilingual benchmarks demonstrate that MeTA-LoRA attains performance on par with—or exceeding—that of full-data LoRA fine-tuning, while using only 1%–5% of the task-specific training data. This substantially improves data efficiency and generalization capability of large language models under constrained-data regimes.

Enables knowledge transfer across tasks with minimal dataImproves data efficiency in multi-task LLM fine-tuningReduces data requirements while maintaining competitive performance

This work addresses the limitations of conventional parameter-efficient fine-tuning (PEFT) methods when applied to sequences of heterogeneous tasks, where shared optimization paths often lead to task interference, negative transfer, and catastrophic forgetting. To mitigate these issues, the study introduces— for the first time—the concept of organizing distinct optimization paths within PEFT, proposing an automated multi-strategy QLoRA framework. This approach automatically groups and orders tasks to construct decoupled, independent adaptation pathways under a fixed parameter budget, enabling efficient co-training across diverse tasks without increasing model size. The method significantly enhances model generalization and transferability, achieving a state-of-the-art performance of 44.78 on the TRACE benchmark and substantially outperforming existing single-strategy PEFT approaches.

Catastrophic ForgettingHeterogeneous TasksOptimization Path

L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models

Feb 07, 2024
HJ
Hyesung Jeon
🏛️ Seoul National University | Sungkyunkwan University

To address the trade-off between accuracy and efficiency in quantization-aware fine-tuning of large language models (LLMs), this paper proposes QAT-LoRA, a tightly integrated framework that unifies quantization-aware training (QAT) with Low-Rank Adaptation (LoRA). It introduces memory-optimized layers supporting 3/4-bit integer quantization and employs layer-wise parameter freezing coupled with gradient reparameterization—enabling high accuracy comparable to fully quantized models while reducing training memory consumption to near-LoRA levels. Experiments on LLaMA and Mistral demonstrate that QAT-LoRA significantly outperforms decoupled approaches such as PTQ+LoRA: under 4-bit quantization, it achieves an average +2.1% improvement in task accuracy and attains few-shot performance nearly matching that of full-precision baselines. The framework thus achieves a unified optimization of low-bit quantization, high accuracy, and low training overhead.

Improves accuracy in low-bit quantization scenariosIntegrates Quantization-Aware Training with LoRA for efficiencyReduces memory and computational costs in large language models

Latest Papers

What's happening recently
View more

This study addresses the inherent trade-off between task-specific performance gains and general capability degradation during large language model fine-tuning by proposing LS-LoRA. The method identifies inter-layer sensitivity variations and introduces input-output cosine similarity as a lightweight forward-pass proxy metric, effectively replacing computationally expensive empirical Fisher information calculations. Based on this metric, LS-LoRA selectively places LoRA adapters in low-sensitivity layers. Experimental results demonstrate that LS-LoRA substantially improves target performance on mathematical and coding tasks while largely preserving general capabilities such as commonsense reasoning. Overall, this work achieves an effective balance between task adaptation and capability preservation under parameter-efficient fine-tuning paradigms.

Capability retentionCatastrophic forgettingLarge language models

This work addresses the trade-off between representational capacity and training efficiency in existing large language model fine-tuning methods, which typically choose between full fine-tuning (FFT) and low-rank adaptation (LoRA). The authors propose MoLF, a novel framework that enables dynamic, continuous mixing of FFT and LoRA at the optimizer level. MoLF employs a gradient-guided routing mechanism to precisely allocate parameter updates across both paths and integrates variable-rank LoRA pairs with a frozen backbone strategy, ensuring accurate gradient signals for both routes while maintaining memory efficiency. Experiments demonstrate that MoLF matches or nearly matches the best performance among FFT and LoRA baselines across all tasks (within ≤1.5% gap). Its efficient variant, MoLF-Efficient, surpasses current adaptive LoRA approaches by up to 20% on the Fact task and achieves gains of up to 9% on Med and SQL tasks.

Fine-TuningFull Fine-TuningLarge Language Models

This work addresses the lack of theoretical understanding regarding the generalization gap between Low-Rank Adaptation (LoRA) and full fine-tuning. Within a simplified linear regression framework, the study systematically compares their generalization behaviors by integrating statistical learning theory with excess risk analysis. It establishes, for the first time, a theoretical guarantee that when the discrepancy between the pre-trained model and the downstream task exhibits a low-rank structure, LoRA achieves lower excess risk than full fine-tuning in both over-parameterized and under-parameterized regimes. Empirical experiments further corroborate this finding, demonstrating a non-intuitive improvement in test accuracy under low-rank constraints and thereby revealing the intrinsic mechanism underlying LoRA’s superior generalization capability.

excess riskfine-tuninggeneralization

This work addresses the sensitivity of LoRA fine-tuning to hyperparameters and the inefficiencies in concurrent multi-task training, which often lead to computational waste and low GPU utilization. The authors propose a cooperative training system that executes multiple LoRA tasks concurrently on a shared frozen backbone. By dynamically monitoring loss trajectories, the system enables early termination of underperforming configurations. It further integrates fused-group GEMM operations with rank-local adapter parallelism to enable joint scheduling both across and within tasks. This approach is the first to exploit synergistic optimization opportunities among multiple LoRA tasks, achieving up to a 13.8× speedup over state-of-the-art methods without compromising adapter performance, thereby substantially improving cluster resource efficiency.

GPU utilizationheterogeneous workloadshyperparameter tuning

This work addresses the high memory overhead associated with maintaining a library of expert policies in multitask robotic reinforcement learning. It introduces low-rank adaptation (LoRA) into reinforcement learning for the first time, leveraging proximal policy optimization (PPO) to perform parameter-efficient fine-tuning of a shared base policy. The proposed approach constructs a memory-efficient policy library while preserving task success rates without significant degradation. By reducing the memory footprint per policy by 20–160×, it achieves 90–95% storage savings. Consequently, deploying 10–50 distinct policies becomes feasible without requiring memory swapping, substantially enhancing the scalability and practicality of policy libraries in real-world robotic applications.

memory efficiencymulti-task roboticsparameter-efficient fine-tuning

Hot Scholars

DX

Deyi Xiong

Professor, College of Intelligence and Computing, Tianjin University, China
Natural Language ProcessingLarge Language ModelsAI4Science
RT

Radu Timofte

Humboldt Professor for AI and Computer Vision, University of Würzburg
Computer VisionMachine LearningAICompression
JY

Jianwei Yin

Professor of Computer Science and Technology, Zhejiang University
Service ComputingComputer ArchitectureDistributed ComputingAI
FD

Franck Dernoncourt

NLP/ML Researcher. MIT PhD.
Machine LearningNeural NetworksNatural Language Processing
PY

Pin-Yu Chen

Principal Research Scientist, IBM Research AI; MIT-IBM Watson AI Lab; RPI-IBM AIRC
AI SafetyGenerative AITrustworthy Machine LearningAdversarial Machine Learning