Score
Designs, implements, and evaluates procedures to adapt pretrained encoder-based models and classifiers—including encoder-only language models, acoustic encoders, multilingual and multimodal encoders—by fine-tuning weights, representation layers, and hyperparameters (e.g., learning rate and parameter subsets). This work produces task-specific encoders and classifier heads through post-training fine-tuning and optimization of trade-offs such as latency, cost, and predictive performance.
Full-parameter fine-tuning of large language models (LLMs) and vision-language models (VLMs) suffers from prohibitive computational costs, overfitting, and catastrophic forgetting. Method: We propose the first structured taxonomy of parameter-efficient fine-tuning (PEFT), encompassing additive, selective, reparameterized, hybrid, and unified frameworks, and conduct the first standardized cross-modal (language/vision) and cross-task (understanding/generation) evaluation. Contribution/Results: Through theoretical analysis and multi-domain transfer experiments, we comprehensively benchmark mainstream PEFT methods—including LoRA, Adapter, and Prompt Tuning—demonstrating up to 95% GPU memory reduction and substantial computational savings while retaining over 90% of full fine-tuning performance. We further uncover fundamental trade-offs among robustness, scalability, and interpretability across PEFT paradigms, establishing a methodological foundation and empirical basis for efficient, reliable, and generalizable multimodal model adaptation.
This work investigates the hierarchical representation evolution of pretrained Transformers (e.g., BERT) during fine-tuning—specifically, how models balance preservation of general-purpose features against acquisition of task-specific ones. We propose an activation interpretability framework based on sparse autoencoders (SAEs), integrating inter-layer similarity metrics, token-level activation visualization, and cross-dataset comparative experiments. Our analysis systematically reveals, for the first time, a universal layer-wise functional specialization pattern: early layers retain generic linguistic features, middle layers exhibit progressive transitional behavior, and deeper layers specialize in task-adaptive representations. This finding provides both theoretical grounding and empirical evidence for controllable fine-tuning, model compression, and interpretable AI.
This study addresses the insufficient exploration of relationships and trade-offs among different paradigms for cross-task adaptation of large language models by constructing the first unified analytical framework. Methodologically, it systematically integrates parameter-efficient fine-tuning, in-context learning, and embedding injection techniques, establishing a comprehensive taxonomy along the dimensions of weights, prompts, and embeddings. Employing a systematic review methodology, the work thoroughly analyzes the strengths, limitations, and intrinsic connections of each paradigm. The findings reveal the evolutionary logic underlying these three paradigms, yielding a holistic taxonomic landscape that clarifies their core advantages and constraints while identifying key open problems. Ultimately, this research provides strategic guidance for future investigations into the efficient adaptation of large models.
This work addresses the problem of efficiently selecting an optimal subset of auxiliary tasks from multiple sources to enhance fine-tuning performance on a target task. We propose a lightweight, training-free subset selection method grounded in meta-initialization. Our core contribution is a first-order gradient approximation framework that leverages first-order Taylor expansion and gradient sensitivity analysis to estimate fine-tuning loss for arbitrary task subsets—entirely on CPU, within seconds. Unlike conventional enumeration or reinforcement learning–based approaches, our method eliminates the need for repeated fine-tuning, achieving a 30× speedup over baselines with only 1% estimation error. Empirically, on instruction tuning and chain-of-thought tuning benchmarks, subsets selected by our method yield up to a 3.8% absolute improvement in downstream task performance, significantly outperforming existing subset selection techniques.
This work addresses the lack of systematic guidance in existing parameter-efficient fine-tuning (PEFT) methods regarding which layers of large language models should be adapted, a gap that complicates the trade-off among performance, training cost, and inference latency. The authors propose a unified projected residual perspective, revealing through local quadratic approximation that inter-layer adaptability is governed by three factors: residual norm, activation energy, and layer coupling. Leveraging this insight, they develop a reusable diagnostic tool—Layer Card—that enables goal-directed, flexible layer selection. This framework provides the first theoretical foundation for layer selection in PEFT, demonstrating on Qwen3-8B that fine-tuning only a subset of layers can closely match the performance of full-layer LoRA while substantially reducing both training overhead and the number of active adapters during inference.
To address the limited cross-lingual transferability and severe domain mismatch of self-supervised learning (SSL) pre-trained models in automatic speech recognition (ASR) for low-resource languages, this paper proposes a lightweight adapter method with *intermediate warm-start*. Under frozen SSL backbone constraints, only 1–5% of parameters are fine-tuned. A two-stage progressive adaptation jointly optimizes adapter architecture and downstream model initialization. The novel intermediate warm-start mechanism mitigates speech feature distribution shift, substantially improving generalization to unseen languages. Evaluated on the ML-SUPERB benchmark, our approach achieves up to 28% relative reduction in character/phone error rates over standard efficient fine-tuning, significantly alleviating the bottleneck in low-resource cross-lingual ASR adaptation.
This work addresses the challenge of selecting effective fine-tuning strategies for encoder-decoder pre-trained language models in generation and question-answering tasks. It proposes the Match Task to Objective (MTO) framework, which establishes the first systematic alignment mechanism between downstream tasks and pre-training objectives. MTO automatically constructs training data and prompt templates that are consistent with the original pre-training objective and extends this alignment to soft prompt tuning, thereby enabling precise task–objective matching. Experimental results demonstrate that MTO achieves over 120% performance improvement under few-shot settings compared to existing methods, significantly outperforms strong baselines in full-data scenarios, and substantially enhances the effectiveness of prompt tuning.
Pretraining for low-resource languages is often constrained by data scarcity, necessitating repeated use of limited corpora and thereby compromising model generalization. This work systematically compares two strategies to mitigate this issue: mixed training with high-resource auxiliary languages and extensive hyperparameter tuning. Through large-scale hyperparameter search (approximately 1,000 runs) combined with μP-based transfer tuning, we demonstrate that mixed training substantially outperforms fine-tuned monolingual baselines—a gap that widens with increasing model scale from 150M to 1.43B parameters. Mixed training yields validation loss improvements equivalent to a 2–3× increase in target-language data and downstream accuracy gains equivalent to a 2–13× data boost. Moreover, we show that validation loss measured solely on the target language systematically underestimates the cross-lingual knowledge benefits conferred by mixed training.
This study systematically investigates whether representation collapse during multi-stage post-training of large language models leads to degraded adaptability, weakened out-of-distribution generalization, and deteriorated calibration. To this end, the authors construct a comprehensive measurement framework encompassing hidden states, logits, token trajectories, and LoRA updates, thereby establishing— for the first time—the causal relationship between representation collapse and declines in model plasticity, generalization, and calibration. Building on these insights, they propose lightweight intervention strategies, including mixed-domain replay, feature refreshing, representation diversity regularization, and decorrelation of LoRA updates. These methods significantly enhance continual learning capability while preserving behavioral gains, effectively mitigating representation collapse.
This study addresses the inherent trade-off between task-specific performance gains and general capability degradation during large language model fine-tuning by proposing LS-LoRA. The method identifies inter-layer sensitivity variations and introduces input-output cosine similarity as a lightweight forward-pass proxy metric, effectively replacing computationally expensive empirical Fisher information calculations. Based on this metric, LS-LoRA selectively places LoRA adapters in low-sensitivity layers. Experimental results demonstrate that LS-LoRA substantially improves target performance on mathematical and coding tasks while largely preserving general capabilities such as commonsense reasoning. Overall, this work achieves an effective balance between task adaptation and capability preservation under parameter-efficient fine-tuning paradigms.
This work addresses the inefficiency of full-model fine-tuning in monolingual language model development for low-resource languages, which incurs high computational costs and fails to leverage modular adaptation effectively. To overcome this limitation, the authors propose an efficient transfer learning approach that integrates a target-language-specific tokenizer, freezes the corresponding embedding layer, and fine-tunes only the remaining model parameters—departing from conventional full-parameter fine-tuning paradigms. Evaluated on Scottish Gaelic, Irish, and Quechua (with only 8.5k training samples), the method consistently outperforms baseline approaches across masked language modeling, named entity recognition, and part-of-speech tagging tasks, demonstrating its effectiveness and generalizability under extremely low-resource conditions.