Score
Designs, implements, and evaluates methods for reusing pretrained models on new tasks or target domains by selecting and applying adaptation strategies such as fine‑tuning, feature extraction with frozen layers, adapter modules, and low‑rank adaptation. Builds transfer-learning and domain-adaptation pipelines, runs experiments and assessments to measure performance, compute and data-efficiency trade-offs, and domain generalization, and analyzes protocols and evaluation metrics to guide method selection.
This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.
This study addresses the **reliable assessment of knowledge transferability** in transfer learning—a longstanding challenge hindered by inconsistent evaluation criteria, poor interpretability, and ill-defined applicability scopes. We propose the first **two-dimensional classification framework**, systematically organizing over 60 mainstream transferability metrics along axes of *transferable knowledge type* (e.g., features, relations, semantics) and *measurement granularity* (sample-, task-, or domain-level), while rigorously reconstructing their mathematical foundations, underlying assumptions, and failure boundaries. Through cross-modal and cross-task empirical analysis, we characterize the efficacy gradients and root limitations of metrics across paradigms (e.g., pretraining-finetuning). Our work establishes a standardized assessment pathway and principled metric selection guidelines for transferability evaluation, advancing trustworthy AI evaluation infrastructure, and identifying key future directions—including dynamic transferability modeling and causally grounded metrics.
Domain adaptation (DA) faces practical challenges including difficulty in problem identification and lack of principled guidance for method selection. Method: This paper proposes the first problem-oriented DA framework, introducing a novel five-dimensional scenario taxonomy that systematically characterizes the causes and patterns of data distribution shift, complemented by a scenario identification guide and a method recommendation mechanism. Grounded in design science research, the framework undergoes iterative empirical evaluation across synthetic and real-world datasets, as well as a 100-participant user study. Contribution/Results: Results demonstrate significant improvements in interpretability, generality, and usability. The framework substantially enhances non-expert users’ accuracy in understanding DA tasks and their rationality in method selection, thereby addressing a critical gap in problem-driven adaptive decision-making research.
This work addresses the challenge of selecting appropriate source domains and pre-trained models for unsupervised domain adaptation when target-domain labels are unavailable—a critical yet underexplored problem that often limits adaptation performance. To tackle this, the authors propose the PAS (Pre-trained Adaptation Suitability) scoring mechanism, which evaluates source–target domain compatibility and model transferability by analyzing the geometry of pre-trained feature embedding spaces. Remarkably, PAS accurately predicts post-adaptation target accuracy without requiring any target labels. This approach enables, for the first time, joint unsupervised selection of both source domains and pre-trained models. Extensive experiments on multiple image classification benchmarks demonstrate a strong correlation between PAS scores and actual adaptation accuracy, leading to significantly improved performance while substantially reducing computational overhead.
This work addresses unsupervised domain adaptation (UDA) for image classification, where labeled source-domain data and unlabeled target-domain data are available. We systematically evaluate mainstream UDA methods on standard benchmarks—Office-31 and Office-Home—within a unified experimental framework. Our key contribution is the first comparative analysis of Transformer-based UDA algorithms (e.g., SSRT) under varying data scales and domain shifts, revealing their robustness boundaries and failure modes. Implementing adversarial training, feature alignment, self-training, and safe self-refinement (SSRT) in PyTorch, we enable reproducible large-scale ablation studies. Results show SSRT achieves 91.6% accuracy on Office-31; however, it suffers significant degradation under small-batch settings—dropping to 72.4% on Office-Home—highlighting a critical practical limitation. This empirical finding provides essential guidance for deploying UDA methods in real-world scenarios with constrained computational resources.
This study addresses the challenge of selecting suitable ImageNet-pretrained models for image classification tasks in target domains by introducing a multidimensional evaluation framework. The authors systematically fine-tune the output layers and general parameters of eleven pretrained models across five diverse datasets, evaluating their performance under both single-run and multiple-run training settings. Through comprehensive assessment of accuracy, accuracy density, training time, and model size, the work quantifies the cross-domain transferability differences among pretrained models, revealing consistent patterns in how model characteristics align with task-specific requirements. These findings provide empirical evidence and practical guidelines for informed model selection in real-world applications.
To address the insufficient generalization of pretrained models when fine-tuned on few-shot, high-dimensional binary classification tasks, this paper proposes a novel fine-tuning framework based on weight matrix reparameterization. The method couples low-rank adaptation (LoRA) with a base-model rescaling mechanism and employs random matrix theory to model the generalization behavior of high-dimensional classifiers, thereby revealing how rescaling governs spectral distribution and generalization bounds. Theoretically, the approach substantially mitigates overfitting by controlling the effective rank and condition number of the classifier’s weight matrix. Empirically, it consistently improves performance across multiple binary classification benchmarks and large language model (LLM) fine-tuning tasks, demonstrating particularly pronounced generalization gains under extreme data scarcity.
In few-shot domain adaptation, two key challenges hinder performance: (1) hyperparameter tuning is impractical due to scarce validation data, and (2) models lack robustness under distribution shift. To address these, we propose Soup-Adapter—a hyperparameter-agnostic, distributionally robust adapter ensemble method. Our approach extends the CLIP adapter paradigm to DINOv2 for the first time and introduces a reparameterizable adapter soup: multiple adapters are trained independently via distinct paths, their outputs are averaged at inference, and their parameters are concatenated and reparameterized to stabilize optimization. Experiments demonstrate that Soup-Adapter consistently outperforms single-adapter baselines across a wide range of hyperparameters and significantly improves both accuracy and robustness on cross-domain few-shot tasks. This work establishes a new, efficient, practical, and generalizable paradigm for few-shot domain adaptation.
This paper addresses the reproducibility challenge of adaptive data selection strategies in transfer learning, exposing a fundamental trade-off between adaptation efficacy and result consistency under dynamic sample prioritization. We formally define selection sensitivity Δ_Q and theoretically prove that the probability of reproducibility failure grows quadratically with Δ_Q but decays exponentially with sample size; furthermore, source-domain pretraining substantially mitigates this risk. Empirical validation on MultiNLI—spanning six mainstream strategies (e.g., gradient-based selection, curriculum learning)—confirms the theory: highly adaptive methods improve performance yet incur >25% failure rates, whereas low-adaptivity strategies maintain <7% failure; source pretraining further reduces failure rates by up to 30%. Our core contribution is the first quantitative, empirically verifiable analytical framework for assessing the reliability of adaptive data selection in transfer learning.
Full-parameter fine-tuning suffers from overfitting and high computational overhead. To address this, we propose BioTune—a selective fine-tuning method that adaptively identifies and updates only the most critical network layers for a target domain, guided by an evolutionary algorithm. Its core innovation lies in dynamically optimizing both parameter efficiency and generalization via differentiable evolutionary search, seamlessly integrating deep transfer learning without relying on manually predefined layer-selection heuristics. Evaluated across nine cross-domain image classification benchmarks, BioTune matches or exceeds the accuracy of state-of-the-art methods—including AutoRGN and LoRA—while drastically reducing trainable parameters by an average of 72.4%. This demonstrates BioTune’s dual advantage: superior parameter efficiency without sacrificing predictive performance.
To address severe catastrophic forgetting and low data efficiency in task adaptation of large language models (LLMs), this paper proposes Continuous Low-Rank Adaptation (CLoRA)—the first framework integrating Low-Rank Adaptation (LoRA) with continual learning. CLoRA introduces a knowledge retention module to mitigate forgetting and an adaptive parameter update strategy to enhance multi-task stability. Further augmented with knowledge distillation, it achieves both model compression and strong generalization in privacy-sensitive settings. Extensive experiments across 15 heterogeneous datasets demonstrate that CLoRA outperforms state-of-the-art baselines by +2.7%–5.3% in average task accuracy, reduces GPU memory consumption by 38%, and cuts training data requirements by 42%. The framework thus delivers superior efficiency, robustness, and practicality for continual LLM adaptation.