Score
Designs and implements the process of adapting a pre‑trained model to a target task or dataset by continuing training on task‑specific data and tuning training choices (learning rates, batch sizes, loss terms, regularization, parameter freezing, etc.). Builds and evaluates the resulting fine‑tuned models, diagnosing issues such as overfitting or catastrophic forgetting, selecting checkpoints, and preparing the adapted model for downstream use or deployment.
Existing theoretical frameworks struggle to explain why larger-scale pre-trained models substantially reduce sample complexity on downstream tasks. This work proposes a novel theoretical framework—termed “caulking”—inspired by parameter-efficient fine-tuning methods such as adapters, low-rank adaptation, and partial fine-tuning. It establishes, for the first time, a provable relationship between the scale of pre-trained models and the sample complexity of downstream tasks. By rigorously linking stronger pre-training capabilities to reduced data requirements in transfer learning, this study not only addresses a critical gap in current theoretical understanding but also provides a solid foundation for empirically observed scaling laws, demonstrating that enhanced pre-training capacity can significantly decrease the number of samples needed for effective downstream adaptation.
This work addresses catastrophic forgetting in fine-tuning pretrained models, where newly acquired knowledge overwrites previously learned information. To mitigate this issue, the authors propose a function-preserving model expansion approach that mathematically duplicates and scales parameters of selected Transformer submodules during initialization. This technique enables stable training and faithful retention of original model capabilities without altering the initial functionality. By circumventing the traditional trade-off between plasticity and stability, the method achieves performance comparable to full fine-tuning while expanding only a minimal number of layers. Consequently, it fully preserves the model’s original knowledge and substantially reduces computational overhead.
In large-scale pretraining, learning rate scheduling critically influences both training efficiency and model performance. This work proposes two paradigms—Fitting and Transfer. The Fitting paradigm establishes, for the first time, a scaling law for learning rate search factors, reducing hyperparameter tuning complexity from O(n³) to O(n·C_D·C_η). The Transfer paradigm extends μTransfer to Mixture-of-Experts (MoE) architectures and generalizes it across multiple hyperparameter dimensions, including depth, weight decay, and token length. Empirical results demonstrate that while μTransfer exhibits limited scalability in large-scale settings, the Fitting paradigm—grounded in the derived scaling law—offers superior scalability and practicality, providing a systematic guideline for hyperparameter tuning in industrial-scale pretraining.
Standard fine-tuning of foundation models suffers from low downstream adaptation efficiency and fails to recover the optimal adaptable parameter set. Method: We propose the first PEFT co-optimization framework that explicitly integrates meta-learning (MAML-style) into the foundation model’s retraining phase, using a LoRA-inspired low-rank adaptation structure. Contribution/Results: We theoretically prove that standard retraining is inherently suboptimal in adaptability, whereas our method strictly recovers the optimal adaptable parameters and provides a generalization error bound. Experiments on RoBERTa with the ConvAI2 dialogue continuation task demonstrate significant improvements in zero-shot and few-shot rapid adaptation performance, empirically validating the theoretical guidance.
This paper addresses the challenge of downstream adaptation for large language models when original training data is inaccessible. We propose a **non-parametric knowledge transfer paradigm**: rather than updating model weights, our approach extracts structured cognitive strategies from teacher model outputs via inference trajectory distillation and implicit behavioral modeling. The method comprises four core components: trajectory contrastive learning, latent state-space alignment, logical formalization distillation, and backpropagation-free policy imitation—constituting the first zero-gradient, memory-efficient knowledge absorption framework. Evaluated on six cross-task generalization benchmarks, our method achieves an average accuracy improvement of 9.2% and reduces inference latency by 37%, significantly outperforming parameter-efficient fine-tuning baselines (e.g., LoRA, QLoRA) and prompt engineering approaches.
This work addresses the critical challenge in continual learning of mitigating catastrophic forgetting during downstream fine-tuning while preserving capabilities acquired during upstream training. The authors propose treating “robustness to subsequent fine-tuning” as a first-class objective in upstream training and systematically investigate data scheduling strategies across a three-stage pipeline—pretraining, post-training, and downstream fine-tuning. Their key finding is that early exposure to post-training data during pretraining—termed “early data exposure”—consistently outperforms pure post-training or conventional mixing strategies, yielding superior trade-offs between upstream knowledge retention and downstream task performance across model scales from 135M to 1B parameters. This approach complements regularization techniques such as replay and Dropout and, under fixed compute budgets, reveals an optimal data allocation scheme.
This work addresses the critical challenge of dynamically determining when to perform continual fine-tuning of foundation models on resource-constrained devices under limited computational budgets to maximize performance. The problem is formally cast, for the first time, as a constrained Markov decision process, where the state encompasses model performance, remaining compute budget, and the relevance of incoming data to the historical distribution. The authors propose an online decision-making strategy based on an Actor-Critic reinforcement learning framework; when fine-tuning gains are predictable, dynamic programming is also integrated for optimal scheduling. Experimental results demonstrate that the proposed approach improves accuracy by over 4% compared to strong baselines under identical budgets and achieves 97% of the performance of full-parameter fine-tuning using only 25% of the fine-tuning steps.
This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.
This work proposes GIFT, a novel framework that actively leverages the confidence signals from instruction-tuned models to guide low-rank adapter training during task-specific fine-tuning—addressing the limitation of existing approaches that treat instruction models merely as passive targets for merging. By incorporating confidence-guided fine-tuning, low-rank adaptation, and model merging, GIFT effectively balances task specialization with general instruction-following capabilities. The optimized adapters are merged back into the original instruction model, preserving its broad competence while enhancing performance on specific tasks. Evaluated across multiple mathematical and knowledge-intensive benchmarks, GIFT significantly outperforms standard fine-tuning and prevailing transfer learning methods, demonstrating superior generalization and test-time scalability.