Score
Design, implement, and evaluate methods to adapt pretrained large or small language model weights and outputs for specific tasks or behaviors, including full-weight fine-tuning, instruction-tuning, adding and training task-specific heads (e.g., regression or classification), transfer learning, and reinforcement-learning-based fine-tuning. Build and apply parameter-efficient and quantization-aware approaches (LoRA, quantized LoRA/QLoRA, low-rank adapters), and perform multilingual or cross-lingual tuning on translated or language-balanced datasets while measuring task metrics, representational drift, adversarial robustness/compliance, and compute/memory trade-offs.
This study addresses the insufficient exploration of relationships and trade-offs among different paradigms for cross-task adaptation of large language models by constructing the first unified analytical framework. Methodologically, it systematically integrates parameter-efficient fine-tuning, in-context learning, and embedding injection techniques, establishing a comprehensive taxonomy along the dimensions of weights, prompts, and embeddings. Employing a systematic review methodology, the work thoroughly analyzes the strengths, limitations, and intrinsic connections of each paradigm. The findings reveal the evolutionary logic underlying these three paradigms, yielding a holistic taxonomic landscape that clarifies their core advantages and constraints while identifying key open problems. Ultimately, this research provides strategic guidance for future investigations into the efficient adaptation of large models.
To address the inefficiency of fine-tuning large language models (LLMs) under resource constraints and their poor adaptability to multimodal tasks, this paper proposes Dynamic LoRA—a lightweight, input-aware fine-tuning method. Its core innovations are a layer-wise importance reallocation mechanism and a differentiable adapter architecture based on low-rank decomposition. By leveraging gradient-sensitive scoring and input-feature-distribution-guided parameter reweighting, Dynamic LoRA unifies task-specific customization with cross-task generalization. Compared to static LoRA, it achieves 88.1% accuracy and 87.3% F1 score on the GLUE benchmark while incurring only a 0.1% increase in computational overhead—significantly improving the efficiency–performance trade-off. This work establishes a novel paradigm for resource-sensitive, multimodal LLM adaptation.
Quantized LLM fine-tuning suffers from “balance mismatch”: LoRA adapters exhibit high input/output complexity yet low effective trainability, leading to underfitting; subsequent low-precision conversion further degrades performance. Method: We propose Q-BaRA and QA-HiRA—Q-BaRA enhances adapter effectiveness via structural simplification and joint rank optimization; QA-HiRA introduces a single-matrix high-rank adapter coupled with block-wise quantization alignment, enabling lossless integration of fine-tuned parameters into 4/8-bit inference models. Contribution/Results: This work is the first to identify and characterize this imbalance, establishing a new quantization-aware end-to-end fine-tuning paradigm. Evaluated on LLaMA and LLaMA2, our methods significantly outperform state-of-the-art approaches in accuracy, while preserving parameter count and computational cost during training and incurring zero additional deployment overhead.
LoRA achieves efficiency in single-task fine-tuning but struggles with cross-task knowledge transfer and relies heavily on task-specific data in multi-task settings. To address this, we propose MeTA-LoRA, a two-stage low-rank adaptation framework: (1) a task-specific stage, where lightweight adapters are learned from minimal per-task samples; and (2) a meta-adaptation stage, where a shared adapter is updated via gradient aggregation across tasks, enabling implicit knowledge transfer. To our knowledge, this is the first work to systematically integrate multi-task knowledge transfer into the LoRA paradigm. Experiments across multi-task and multilingual benchmarks demonstrate that MeTA-LoRA attains performance on par with—or exceeding—that of full-data LoRA fine-tuning, while using only 1%–5% of the task-specific training data. This substantially improves data efficiency and generalization capability of large language models under constrained-data regimes.
This work addresses the limitations of conventional parameter-efficient fine-tuning (PEFT) methods when applied to sequences of heterogeneous tasks, where shared optimization paths often lead to task interference, negative transfer, and catastrophic forgetting. To mitigate these issues, the study introduces— for the first time—the concept of organizing distinct optimization paths within PEFT, proposing an automated multi-strategy QLoRA framework. This approach automatically groups and orders tasks to construct decoupled, independent adaptation pathways under a fixed parameter budget, enabling efficient co-training across diverse tasks without increasing model size. The method significantly enhances model generalization and transferability, achieving a state-of-the-art performance of 44.78 on the TRACE benchmark and substantially outperforming existing single-strategy PEFT approaches.
To address the trade-off between accuracy and efficiency in quantization-aware fine-tuning of large language models (LLMs), this paper proposes QAT-LoRA, a tightly integrated framework that unifies quantization-aware training (QAT) with Low-Rank Adaptation (LoRA). It introduces memory-optimized layers supporting 3/4-bit integer quantization and employs layer-wise parameter freezing coupled with gradient reparameterization—enabling high accuracy comparable to fully quantized models while reducing training memory consumption to near-LoRA levels. Experiments on LLaMA and Mistral demonstrate that QAT-LoRA significantly outperforms decoupled approaches such as PTQ+LoRA: under 4-bit quantization, it achieves an average +2.1% improvement in task accuracy and attains few-shot performance nearly matching that of full-precision baselines. The framework thus achieves a unified optimization of low-bit quantization, high accuracy, and low training overhead.
This study addresses the inherent trade-off between task-specific performance gains and general capability degradation during large language model fine-tuning by proposing LS-LoRA. The method identifies inter-layer sensitivity variations and introduces input-output cosine similarity as a lightweight forward-pass proxy metric, effectively replacing computationally expensive empirical Fisher information calculations. Based on this metric, LS-LoRA selectively places LoRA adapters in low-sensitivity layers. Experimental results demonstrate that LS-LoRA substantially improves target performance on mathematical and coding tasks while largely preserving general capabilities such as commonsense reasoning. Overall, this work achieves an effective balance between task adaptation and capability preservation under parameter-efficient fine-tuning paradigms.
This work addresses the trade-off between representational capacity and training efficiency in existing large language model fine-tuning methods, which typically choose between full fine-tuning (FFT) and low-rank adaptation (LoRA). The authors propose MoLF, a novel framework that enables dynamic, continuous mixing of FFT and LoRA at the optimizer level. MoLF employs a gradient-guided routing mechanism to precisely allocate parameter updates across both paths and integrates variable-rank LoRA pairs with a frozen backbone strategy, ensuring accurate gradient signals for both routes while maintaining memory efficiency. Experiments demonstrate that MoLF matches or nearly matches the best performance among FFT and LoRA baselines across all tasks (within ≤1.5% gap). Its efficient variant, MoLF-Efficient, surpasses current adaptive LoRA approaches by up to 20% on the Fact task and achieves gains of up to 9% on Med and SQL tasks.
This work addresses the lack of theoretical understanding regarding the generalization gap between Low-Rank Adaptation (LoRA) and full fine-tuning. Within a simplified linear regression framework, the study systematically compares their generalization behaviors by integrating statistical learning theory with excess risk analysis. It establishes, for the first time, a theoretical guarantee that when the discrepancy between the pre-trained model and the downstream task exhibits a low-rank structure, LoRA achieves lower excess risk than full fine-tuning in both over-parameterized and under-parameterized regimes. Empirical experiments further corroborate this finding, demonstrating a non-intuitive improvement in test accuracy under low-rank constraints and thereby revealing the intrinsic mechanism underlying LoRA’s superior generalization capability.
This work addresses the sensitivity of LoRA fine-tuning to hyperparameters and the inefficiencies in concurrent multi-task training, which often lead to computational waste and low GPU utilization. The authors propose a cooperative training system that executes multiple LoRA tasks concurrently on a shared frozen backbone. By dynamically monitoring loss trajectories, the system enables early termination of underperforming configurations. It further integrates fused-group GEMM operations with rank-local adapter parallelism to enable joint scheduling both across and within tasks. This approach is the first to exploit synergistic optimization opportunities among multiple LoRA tasks, achieving up to a 13.8× speedup over state-of-the-art methods without compromising adapter performance, thereby substantially improving cluster resource efficiency.
This work addresses the high memory overhead associated with maintaining a library of expert policies in multitask robotic reinforcement learning. It introduces low-rank adaptation (LoRA) into reinforcement learning for the first time, leveraging proximal policy optimization (PPO) to perform parameter-efficient fine-tuning of a shared base policy. The proposed approach constructs a memory-efficient policy library while preserving task success rates without significant degradation. By reducing the memory footprint per policy by 20–160×, it achieves 90–95% storage savings. Consequently, deploying 10–50 distinct policies becomes feasible without requiring memory swapping, substantially enhancing the scalability and practicality of policy libraries in real-world robotic applications.