Score
Designs and implements workflows, training pipelines, and procedures to adapt pre‑trained transformer/foundation models (including LLMs) to downstream tasks via supervised and task‑specific fine‑tuning. This includes developing and evaluating parameter‑efficient techniques (e.g., LoRA, adapters), supervised fine‑tuning (SFT), hyperparameter and autotuning frameworks, and tooling to optimize, validate, and prepare fine‑tuned large models for deployment.
Parameter-efficient fine-tuning (PEFT) methods for large foundation models suffer from unclear cross-modal adaptation mechanisms and a lack of systematic classification frameworks. Method: This paper introduces the first unified PEFT taxonomy spanning language, vision, and multimodal foundation models, systematically analyzing the adaptation principles and architectural distinctions of mainstream techniques—including LoRA, Adapter, Prompt Tuning, and IA³—within the Transformer architecture. Through comparative analysis, it identifies evolutionary trends and critical research gaps. Contribution/Results: We publicly release an authoritative, curated PEFT literature repository on GitHub, featuring a structured knowledge graph and practical implementation guidelines. This resource significantly lowers the barrier to task-specific customization of large models and establishes foundational support for advancing PEFT theory and enabling robust cross-modal applications.
研究探讨了客户服务LLM的多任务处理策略,通过多任务微调、顺序更新或模型合并的方法,发现多任务全微调在所有测试模型尺寸中表现最佳。
Standard fine-tuning of foundation models suffers from low downstream adaptation efficiency and fails to recover the optimal adaptable parameter set. Method: We propose the first PEFT co-optimization framework that explicitly integrates meta-learning (MAML-style) into the foundation model’s retraining phase, using a LoRA-inspired low-rank adaptation structure. Contribution/Results: We theoretically prove that standard retraining is inherently suboptimal in adaptability, whereas our method strictly recovers the optimal adaptable parameters and provides a generalization error bound. Experiments on RoBERTa with the ConvAI2 dialogue continuation task demonstrate significant improvements in zero-shot and few-shot rapid adaptation performance, empirically validating the theoretical guidance.
This study systematically evaluates the applicability of the foundation model (FM) paradigm across three scientific domains—genomics, satellite imagery, and time-series analysis—to assess whether FMs can supplant traditional supervised learning. Method: We construct a cross-modal benchmark framework employing lightweight architectures (e.g., Wide ResNet, U-Net), automated hyperparameter optimization, and standardized training protocols to rigorously compare domain-specific FMs against strong supervised baselines. Contribution/Results: Across all tasks, carefully tuned supervised models match or exceed state-of-the-art domain-specific FMs; large-scale pretraining yields no consistent empirical gains. This work provides the first multi-modal scientific validation that the FM paradigm remains immature for these domains. We open-source two automated evaluation workflows and underscore the necessity—and benchmarking value—of strong supervised baselines in scientific AI assessment.
This study addresses the prohibitively high computational cost of full-parameter fine-tuning for large language models (LLMs) in unit test generation—a key barrier to practical deployment. We conduct the first systematic empirical evaluation of three parameter-efficient fine-tuning (PEFT) methods—LoRA, (IA)³, and Prompt Tuning—across multiple LLM scales. On standard code testing benchmarks, all three approaches achieve generation quality comparable to full fine-tuning while updating fewer than 0.5% of model parameters. Prompt Tuning delivers the best trade-off between resource efficiency and effectiveness, yielding the lowest overall cost; LoRA most closely matches full fine-tuning performance across diverse scenarios; and (IA)³ exhibits robustness but slower convergence. Our work fills a critical empirical gap in applying PEFT to automated software testing and provides a reproducible methodology and practical guidance for lightweight, LLM-driven test generation.
This work addresses the high computational cost of conventional fine-tuning for instance segmentation, which typically requires updating a large fraction (40–55%) of parameters in large pre-trained models. To improve parameter efficiency, the study explores parameter-efficient fine-tuning (PEFT) methods, introducing LoRA into deformable attention mechanisms for the first time and systematically evaluating the trade-offs between performance and efficiency based on the number and placement of adapters within the Transformer architecture. Experimental results demonstrate that by fine-tuning only 1–6% of the model parameters, the proposed approach matches or even surpasses the performance of full fine-tuning across four benchmark datasets. These findings validate the efficacy and feasibility of PEFT for instance segmentation and further reveal that its effectiveness is influenced by dataset complexity and model architecture.
研究通过系统性调整学习率、批次大小等参数,对比LoRA和全微调方法,在不同模型和数据集上优化监督微调效果。
研究通过控制实验,比较了SFT与LoRA、RL(GRPO)及两者结合的方法在不同规模Qwen3模型上工具调用性能的影响,发现SFT与LoRA在多数情况下表现最佳。
This study addresses the inherent trade-off between task-specific performance gains and general capability degradation during large language model fine-tuning by proposing LS-LoRA. The method identifies inter-layer sensitivity variations and introduces input-output cosine similarity as a lightweight forward-pass proxy metric, effectively replacing computationally expensive empirical Fisher information calculations. Based on this metric, LS-LoRA selectively places LoRA adapters in low-sensitivity layers. Experimental results demonstrate that LS-LoRA substantially improves target performance on mathematical and coding tasks while largely preserving general capabilities such as commonsense reasoning. Overall, this work achieves an effective balance between task adaptation and capability preservation under parameter-efficient fine-tuning paradigms.
This work addresses the challenge of explicitly controlling overfitting during fine-tuning of pretrained Transformers by formulating it as a bilevel optimization regularized framework. It introduces, for the first time, a linear programming–driven local search mechanism that leverages validation gradients and training Hessian information from a warm-up phase to construct a validation-aware descent direction. This enables joint, task-adaptive optimization of both model parameters and regularization hyperparameters without requiring repeated full retraining. Experimental results demonstrate significant reductions in test perplexity on GPT-2 Small and WikiText-2, with particularly pronounced gains in settings prone to overfitting. The approach consistently yields stable improvements across diverse layer configurations and regularization settings.