domain adaptation

Designs, builds, and evaluates methods to transfer and adapt models across differing data distributions by aligning or fusing source and target features, performing data‑efficient or few‑shot fine‑tuning, enabling source‑free/target‑only adaptation, and adjusting priors or parameters to minimize overhead. This work also develops transfer analyses and evaluation protocols to measure cross‑domain generalization and robustness, and includes modality‑specific strategies such as text‑only or simulation‑to‑real adaptation.

domainadaptation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$212K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the unclear interaction between feature alignment and target fitting in cross-modal fine-tuning, which often leads to a mismatch between feature-label structures across source and target domains, thereby degrading generalization. For the first time, this study theoretically characterizes their relationship by introducing the notion of “feature-label distortion,” and establishes a provable generalization bound on target error. Based on this analysis, a principle for joint optimization of alignment and fitting is derived. The resulting framework offers interpretable and actionable design guidelines for cross-modal fine-tuning. Extensive experiments demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across multiple benchmark datasets, confirming its effectiveness and broad applicability.

cross-modal fine-tuningfeature alignmentfeature-label distortion

This work addresses the lack of a unified, rigorous, and realistic evaluation protocol in few-shot transfer learning, which has led to unreliable method comparisons. To this end, we introduce the FEWTRANS benchmark—comprising ten diverse datasets—and the Hyperparameter Ensemble (HPE) evaluation protocol, which effectively mitigates validation set hallucination under data scarcity. Using this framework, we systematically demonstrate for the first time that the choice of pretrained model is more critical than the complexity of the transfer algorithm. We also quantify the performance collapse of multimodal models in specialized domains due to linguistic rarity. Our analysis reveals that simple full-parameter fine-tuning consistently outperforms most sophisticated methods, owing to its ability to flexibly reshape distributed representations and high-level semantic features. The FEWTRANS benchmark is publicly released to provide the community with a reproducible evaluation standard.

benchmarkevaluation protocolfew-shot transfer

Efficient Model Development through Fine-tuning Transfer

Mar 25, 2025
PL
Pin-Jie Lin
🏛️ Virginia Tech | University of Toronto | Vector Institute

Repeated alignment and domain-/language-specific fine-tuning across large language model (LLM) version updates incur prohibitive computational costs. Method: We propose cross-version fine-tuning update transfer: extracting the weight delta vector from a fine-tuned source model, mapping it across versions via a lightweight linear transformation, and directly injecting it into the target base model—bypassing full retraining. Contribution/Results: This work provides the first systematic empirical validation of fine-tuning update transferability across LLM versions. We identify parameter-space linear connectivity as the key prerequisite for effective transfer and introduce a “recovery + fine-tuning” iterative development paradigm. Experiments show absolute accuracy gains of +10.7% on GPQA—surpassing Llama 3.1 8B Instruct; +4.7% and +15.5% on Global MMLU for Malagasy and Turkish, respectively; and substantial reduction in downstream fine-tuning compute overhead.

Avoid repeating alignment for new pretrained model versionsImprove multilingual performance without retraining target modelsTransfer fine-tuning updates between model versions efficiently

Domain Adaptation via Feature Refinement

Aug 22, 2025
SK
Savvas Karatsiolis
🏛️ CYENS Center of Excellence | University of Twente

To address feature inconsistency arising from source-target domain distribution shifts in unsupervised domain adaptation, this paper proposes a lightweight and efficient feature refinement framework. Methodologically, it innovatively integrates batch normalization (BN) statistic adaptation, source-model feature distillation, and hypothesis transfer—requiring neither target-domain labels, complex network architectures, nor auxiliary training objectives. The framework achieves dual alignment: statistical alignment via BN parameter calibration and representational alignment via feature distribution matching. Evaluated on benchmarks including CIFAR10-C, the method substantially outperforms existing state-of-the-art approaches, demonstrating superior robustness to corruptions, higher cross-domain mutual information, and reduced sensitivity to input perturbations. It provides a concise yet effective solution for constructing robust, generalizable domain-invariant feature spaces.

Aligning feature distributions without target labelsImproving model robustness to corruption across domainsUnsupervised domain adaptation under distribution shift

Adaptive Sample Aggregation In Transfer Learning

Aug 29, 2024
SH
Steve Hanneke
🏛️ Purdue University | Columbia University

This work addresses the optimal adaptive aggregation of source and target domain samples in transfer learning to minimize the target risk. To overcome the limitation of existing methods—namely, their inability to uniformly handle diverse distribution divergence measures—we propose a unified weak/strong transfer modulus framework. This is the first approach that automatically adapts to multiple divergence classes—including Wasserstein distance and integral probability metrics (IPMs)—and characterizes their statistical limits. By integrating confidence-set reduction, modulus upper-bound derivation, and adaptive weighted estimation, we achieve near-optimal convergence rates even when the transfer modulus is unknown. Theoretical analysis further reveals that, under causal modeling assumptions, the framework yields provable generalization gains beyond standard transfer bounds. Extensive experiments demonstrate significant improvements in cross-domain classification and regression performance.

Adaptive procedures for strong modulus in complex transfer scenariosUnderstanding statistical limits and achieving faster transfer ratesUnified algorithmic approaches for diverse divergence measures in transfer learning

Latest Papers

What's happening recently
View more

本文提出了一种针对数据可用性随时间演变的领域迁移学习问题(TrED),并探讨了现有方法在处理整个演化过程中的不足。

Data AvailabilityEvaluation CriterionEvolving Domains

This study addresses the high computational cost incurred by repeatedly fine-tuning diffusion models for each target domain in multi-target domain adaptation. To this end, we propose MUSE, a novel framework that introduces a pioneering decoupled adaptation mechanism separating shared semantics from private styles. Through branch-decoupled training, MUSE requires only a single source-guided fine-tuning pass to synthesize domain-specific data for multiple target domains, thereby effectively eliminating redundant fine-tuning overhead. Experimental results demonstrate that MUSE significantly reduces fine-tuning costs while improving average accuracy across target domains in unsupervised domain adaptation tasks, achieving a superior accuracy-efficiency trade-off.

Diffusion ModelsFine-tuning EfficiencyMulti-target Adaptation

This work proposes a multi-source transfer learning method for regression tasks under severe target-domain data scarcity, without requiring assumptions of covariate or label shift. The approach constructs conditional generative models for each heterogeneous source domain and aligns their distributions to the target domain through conditional quantile matching, thereby enabling high-quality data augmentation. Theoretical analysis establishes the convergence rate of the quantile-matching estimator and derives a tighter excess risk bound for the empirical risk minimizer trained on the augmented data. Experimental results demonstrate that the proposed method significantly outperforms both target-only baselines and existing transfer learning approaches on both synthetic and real-world datasets.

data scarcitydomain adaptationheterogeneous sources

This work addresses the challenge of selecting appropriate source domains and pre-trained models for unsupervised domain adaptation when target-domain labels are unavailable—a critical yet underexplored problem that often limits adaptation performance. To tackle this, the authors propose the PAS (Pre-trained Adaptation Suitability) scoring mechanism, which evaluates source–target domain compatibility and model transferability by analyzing the geometry of pre-trained feature embedding spaces. Remarkably, PAS accurately predicts post-adaptation target accuracy without requiring any target labels. This approach enables, for the first time, joint unsupervised selection of both source domains and pre-trained models. Extensive experiments on multiple image classification benchmarks demonstrate a strong correlation between PAS scores and actual adaptation accuracy, leading to significantly improved performance while substantially reducing computational overhead.

domain adaptationpre-trained modelssource selection

Synthetic Dataset Evaluation Based on Generalized Cross Validation

Sep 14, 2025
ZS
Zhihang Song
🏛️ Tsinghua University

Existing synthetic data evaluation lacks unified, transferable quantitative metrics. This paper proposes a novel evaluation framework grounded in generalized cross-validation (GCV) and domain transfer learning. It constructs a cross-dataset performance matrix and defines two core metrics: *fidelity*, quantifying distributional similarity between synthetic and real data; and *generalization coverage*, measuring the task-transfer capability of synthetic data across diverse real-world source domains. The framework is model-agnostic and enables normalized, comparative evaluation of detectors such as YOLOv5s across heterogeneous datasets—including Virtual KITTI, KITTI, and BDD100K. Experiments demonstrate that the method effectively quantifies synthetic data quality, significantly enhancing evaluation generality, comparability, and utility for model optimization. It establishes a scalable, reproducible, and standardized evaluation paradigm for synthetic data development.

Evaluating synthetic dataset quality lacks standard frameworkProposing cross-validation and transfer learning for assessmentQuantifying simulation and transfer quality across domains

Hot Scholars

TL

Tianming Liu

Distinguished Research Professor of Computer Science, University of Georgia
BrainBrain-Inspired AILLMArtificial General Intelligence
NY

Nan Yin

Mohamed bin Zayed University of Artificial Intelligence
Graph Neural NetworksMachine LearningAI4Science
DH

Dongxiao He

Tianjin University, Professor of Computer Science
Graph machine learningNetwork mining
MW

Min Wu

Professor, IEEE Fellow, China University of Geosciences
Process controlRobust controlIntelligent systems
RH

Ruize Han

SUAT
Computer VisionMultimedia AnalysisVideo UnderstandingActive Vision