Score
Designs and implements evaluation protocols, dataset splits, and sampling strategies that measure model performance under domain shift by ensuring non‑overlap of classes/instances, performing bidirectional and cross-dataset tests, and controlling for dataset size; builds comparison and diagnostic analyses that quantify out‑of‑domain transfer, reveal memorization versus true generalization, and report matched‑size cross‑domain performance.
This survey systematically addresses distribution shift between training and deployment in machine learning, focusing on two fundamental challenges: covariate shift (changes in input feature distributions) and concept shift (changes in semantic or class-conditional label distributions). We formalize and unify shift taxonomy, integrating techniques—including distribution shift detection, uncertainty estimation, domain adaptation, anomaly identification, causal inference, and invariant representation learning—within a cohesive framework bridging statistical learning and deep learning. Our key contributions include: (i) a novel robust modeling framework designed to handle heterogeneous shift types; (ii) the first systematic taxonomy covering out-of-distribution (OOD) scenarios; and (iii) a critical analysis revealing limitations of existing methods in jointly mitigating multiple concurrent shifts and generalizing to unseen classes. We establish principled evaluation criteria and outline future research directions—particularly addressing compound shifts and semantic evolution—thereby filling a critical gap in prior surveys, which largely overlook real-world deployment complexities involving intertwined and dynamically evolving shifts.
This study addresses two core challenges in transfer learning: quantitative assessment of knowledge transferability and assurance of trustworthiness. First, it systematically formalizes transfer learning from a trustworthiness perspective, proposing a novel “transferability–trustworthiness” co-evaluation framework; theoretically characterizes transferability bounds under non-IID settings; and develops a new transfer paradigm incorporating multi-dimensional trust constraints—privacy, robustness, and fairness. Methodologically, it integrates statistical learning theory, adversarial robustness analysis, fairness metrics, differential privacy mechanisms, and mainstream transfer approaches (e.g., domain adaptation, meta-transfer learning, federated transfer learning). Key contributions include: (1) establishing the first holistic framework spanning theoretical modeling, quantitative trust attribute measurement, and empirical validation; and (2) identifying three open research directions—trustworthy non-IID transfer, standardized trustworthiness benchmarks, and cross-domain causal generalization.
This study addresses the **reliable assessment of knowledge transferability** in transfer learning—a longstanding challenge hindered by inconsistent evaluation criteria, poor interpretability, and ill-defined applicability scopes. We propose the first **two-dimensional classification framework**, systematically organizing over 60 mainstream transferability metrics along axes of *transferable knowledge type* (e.g., features, relations, semantics) and *measurement granularity* (sample-, task-, or domain-level), while rigorously reconstructing their mathematical foundations, underlying assumptions, and failure boundaries. Through cross-modal and cross-task empirical analysis, we characterize the efficacy gradients and root limitations of metrics across paradigms (e.g., pretraining-finetuning). Our work establishes a standardized assessment pathway and principled metric selection guidelines for transferability evaluation, advancing trustworthy AI evaluation infrastructure, and identifying key future directions—including dynamic transferability modeling and causally grounded metrics.
Existing synthetic data evaluation lacks unified, transferable quantitative metrics. This paper proposes a novel evaluation framework grounded in generalized cross-validation (GCV) and domain transfer learning. It constructs a cross-dataset performance matrix and defines two core metrics: *fidelity*, quantifying distributional similarity between synthetic and real data; and *generalization coverage*, measuring the task-transfer capability of synthetic data across diverse real-world source domains. The framework is model-agnostic and enables normalized, comparative evaluation of detectors such as YOLOv5s across heterogeneous datasets—including Virtual KITTI, KITTI, and BDD100K. Experiments demonstrate that the method effectively quantifies synthetic data quality, significantly enhancing evaluation generality, comparability, and utility for model optimization. It establishes a scalable, reproducible, and standardized evaluation paradigm for synthetic data development.
In transfer learning, blind adaptation from poorly aligned source to target domains often degrades performance. Existing similarity measures rely solely on feature distribution alignment, neglecting label structure and decision boundary relationships, thus failing to reliably predict transfer efficacy. To address this, we propose the Cross-Learning Score (CLS), a dataset similarity metric grounded in bidirectional generalization performance. CLS establishes, for the first time, a theoretical link between dataset similarity and cosine similarity of decision boundaries. It introduces a three-region classification framework—forward, ambiguous, and negative transfer—enabling principled transferability assessment. Compatible with encoder-head architectures, CLS avoids costly high-dimensional distribution estimation and ensures computational efficiency. Extensive experiments across synthetic and real-world benchmarks demonstrate that CLS reliably predicts transfer gain, significantly enhancing the scientific rigor and robustness of source dataset selection.
This work investigates the intrinsic relationships among classifier calibration, confidence quantification, and accuracy prediction under data distribution shift. Theoretically, we establish—for the first time—the computational equivalence of these three tasks, proving their mutual reducibility. Methodologically, we develop a unified cross-task adaptation framework grounded in this reducibility, systematically reusing and enhancing classical algorithms—including Platt scaling, expectation-maximization–based confidence quantification, and accuracy regression. Empirically, our general-purpose approach achieves performance on par with or superior to task-specific methods across multiple distribution-shift benchmarks (e.g., ImageNet-C, CIFAR-10-C). By unifying previously disjoint objectives, this work bridges longstanding task boundaries, offering both a coherent theoretical foundation and a practical, implementation-ready methodology for building distribution-shift-robust predictive models.
This work addresses the fundamental lack of fairness and robustness in evaluating model transferability across domains. We propose the first systematic, standardized benchmarking framework for assessing cross-domain transfer capability. Our method introduces a unified multi-source-domain–target-domain evaluation protocol, encompassing diverse transfer tasks and perturbation-robustness analysis, and adopts head-training (i.e., linear-probe fine-tuning) as the consistent evaluation paradigm. Empirical analysis reveals significant performance discrepancies among existing transferability metrics under varying experimental settings, undermining their reliability. Our framework substantially improves assessment fidelity, yielding an average 3.5% gain in transfer performance under standard head-training configurations. To foster reproducibility and rigorous comparison, we fully open-source all code, datasets, and evaluation pipelines—establishing a new, standardized paradigm for transferability measurement.
To address the challenge of unsupervised model performance estimation under covariate shift—where ground-truth labels are unavailable or delayed post-deployment—this paper proposes the Probability-Adaptive Performance Estimation (PAPE) framework. PAPE requires neither access to true labels nor knowledge of the original model’s architecture or feature representations; it operates solely on the model’s probabilistic outputs and confidence scores. By jointly leveraging density ratio estimation and performance generalization bound theory, PAPE models prediction distributions and applies adaptive reweighting to yield unbiased estimates of arbitrary classification metrics—without assuming a specific shift form or resorting to feature learning or generative modeling. Extensive evaluation across 900+ real-world census dataset–model combinations demonstrates that PAPE reduces mean absolute error by 37% compared to state-of-the-art proxy metrics and drift detection methods, significantly enhancing the reliability and generality of model monitoring in production environments.
This work addresses the challenging problem of model performance estimation under distribution shift when ground-truth labels are unavailable. The authors propose Fused Reference Alignment Prediction (FRAP), a novel approach that synergistically combines the generalization capability of external foundation models with the domain-specific expertise of the target task model. FRAP calibrates prediction distributions via temperature scaling, aligns them by minimizing KL divergence, and generates a robust, domain-adaptive pseudo-label reference distribution through confidence-weighted fusion. Extensive experiments demonstrate that FRAP consistently and substantially outperforms existing performance estimation methods across diverse datasets and model architectures, achieving stable and significant improvements.
This work addresses the challenge of out-of-distribution generalization in geospatial data across regions, where distributional shifts hinder model transfer and existing approaches lack effective means to quantify domain discrepancies. To this end, we propose GeoSpOT, which for the first time integrates optimal transport theory with geographic coordinate encoding to construct a geospatial inter-domain distance metric using only latitude and longitude. Notably, GeoSpOT requires no downstream task data to quantify domain shift or predict model transfer performance. Experimental results demonstrate that the GeoSpOT distance accurately forecasts cross-regional generalization outcomes and effectively guides data selection and identification of high-risk regions.