domain generalization

Designs and evaluates algorithms, datasets, and experimental protocols that measure and improve model performance under distribution shift across domains; builds cross-domain benchmarks, domain sampling strategies, multi-domain experiments, and robustness metrics to assess and compare generalization to unseen environments.

domaingeneralization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Synthetic Dataset Evaluation Based on Generalized Cross Validation

Sep 14, 2025
ZS
Zhihang Song
🏛️ Tsinghua University

Existing synthetic data evaluation lacks unified, transferable quantitative metrics. This paper proposes a novel evaluation framework grounded in generalized cross-validation (GCV) and domain transfer learning. It constructs a cross-dataset performance matrix and defines two core metrics: *fidelity*, quantifying distributional similarity between synthetic and real data; and *generalization coverage*, measuring the task-transfer capability of synthetic data across diverse real-world source domains. The framework is model-agnostic and enables normalized, comparative evaluation of detectors such as YOLOv5s across heterogeneous datasets—including Virtual KITTI, KITTI, and BDD100K. Experiments demonstrate that the method effectively quantifies synthetic data quality, significantly enhancing evaluation generality, comparability, and utility for model optimization. It establishes a scalable, reproducible, and standardized evaluation paradigm for synthetic data development.

Evaluating synthetic dataset quality lacks standard frameworkProposing cross-validation and transfer learning for assessmentQuantifying simulation and transfer quality across domains

Sampling and Loss Weights in Multi-Domain Training

Nov 10, 2025
MS
Mahdi Salmani
🏛️ University of Southern California | Google Research

Training large language models on heterogeneous multi-source data (e.g., Wikipedia, GitHub) suffers from sampling imbalance and loss imbalance, leading to high gradient variance and degraded generalization. Method: This paper systematically analyzes the complementary roles of sampling weights and loss weights in suppressing gradient variance and narrowing the generalization gap. We propose a joint optimization framework grounded in linear regression theory and SGD dynamics, deriving principled co-design criteria for both weight types through theoretical analysis and empirical validation. Contribution/Results: Our key insight is the first formal characterization that sampling and loss weights are not independently tunable but must be jointly configured to simultaneously ensure gradient stability and strong cross-domain generalization. Experiments demonstrate that our method significantly reduces training variance, accelerates convergence, and improves out-of-distribution generalization across diverse domains.

Balancing loss weights to improve generalization performanceOptimizing domain sampling weights for heterogeneous training dataStudying joint dynamics of sampling and loss weights

This work addresses the challenging problem of model performance estimation under distribution shift when ground-truth labels are unavailable. The authors propose Fused Reference Alignment Prediction (FRAP), a novel approach that synergistically combines the generalization capability of external foundation models with the domain-specific expertise of the target task model. FRAP calibrates prediction distributions via temperature scaling, aligns them by minimizing KL divergence, and generates a robust, domain-adaptive pseudo-label reference distribution through confidence-weighted fusion. Extensive experiments demonstrate that FRAP consistently and substantially outperforms existing performance estimation methods across diverse datasets and model architectures, achieving stable and significant improvements.

distribution shiftdomain expertisegeneralization

A General Approach for Determining Applicability Domain of Machine Learning Models

May 28, 2024
LE
Lane E. Schultz
🏛️ University of Wisconsin-Madison | Carnegie Mellon University

This study addresses the lack of generality and interpretability in applicability domain (AD) estimation for machine learning models. We propose a unified AD assessment framework based on kernel density estimation (KDE), which quantifies the distance of a query sample from the training data distribution in feature space and establishes a quantitative relationship among distance, prediction error, and uncertainty. Chemical prior knowledge is incorporated to calibrate the AD decision threshold. To our knowledge, this is the first method enabling consistent, cross-model and cross-task AD evaluation across diverse models—including random forests (RF), gradient-boosted decision trees (GBDT), and graph neural networks (GNN)—and heterogeneous materials datasets (crystals, molecules, alloys). Experiments demonstrate that large KDE-derived distances strongly correlate with high prediction residuals and elevated uncertainty estimates. An open-source toolkit enables automated in-domain/out-of-domain classification. The implementation and documentation are publicly available.

Assessing data distance in feature space for domain determinationDetermining applicability domain of machine learning modelsIdentifying in-domain versus out-of-domain predictions reliably

Latest Papers

What's happening recently
View more

This work addresses the critical challenge of evaluating model generalization in high-stakes scenarios with scarce labels, where existing methods lack reliable, label-free metrics for pre-deployment model selection and post-deployment performance monitoring. To bridge this gap, the study introduces, for the first time, the internal causal circuit mechanisms of Vision Transformers into generalization assessment, proposing two novel unsupervised metrics: Dependency Depth Bias and Circuit Shift Score. The former quantifies depth-wise biases in representational dependency structures, while the latter measures changes in circuit stability under distribution shifts. Extensive experiments across diverse tasks demonstrate that these metrics achieve substantially higher correlations with true generalization performance—improving by 13.4% and 34.1% on average over current approaches—thereby significantly enhancing the reliability of generalization prediction without requiring ground-truth labels.

distribution shiftgeneralizationlabel-free evaluation

Confidence and Dispersity as Signals: Unsupervised Model Evaluation and Ranking

Oct 03, 2025
WD
Weijian Deng
🏛️ The Australian National University | University of Canberra

This paper addresses the problem of evaluating and ranking model generalization under distribution shift in the absence of test-set labels—covering both dataset-centric (evaluating a single model across multiple test sets) and model-centric (ranking multiple models on a single test set) deployment scenarios. We propose a hybrid unsupervised evaluation metric that jointly leverages prediction confidence and inter-class dispersion, and introduce the nuclear norm as a novel, efficient, and robust unified measure computed directly from the model’s output probability distributions. Unlike prior approaches, our method requires no ground-truth labels and imposes no architectural assumptions. Extensive experiments demonstrate that it consistently outperforms confidence-only or dispersion-only baselines across diverse settings—including multi-task learning, various distribution shifts, class imbalance, and real-world datasets—achieving superior generalizability and practical utility.

Developing hybrid metrics for robust unsupervised model performance assessmentEvaluating model generalization without labeled test data under distribution shiftsRanking models and datasets using confidence and dispersity prediction signals

Multi-domain performance analysis with scores tailored to user preferences

Dec 09, 2025
SP
Sébastien Piérard
🏛️ University of Liège

This work addresses two key limitations in multi-domain algorithm performance evaluation: (i) the neglect of user preferences in assessment, and (ii) the masking of domain-specific performance disparities by conventional arithmetic averaging. To this end, we propose a user-preference-parameterized weighted scoring framework. Methodologically, we introduce, for the first time, a continuous family of scoring functions to model performance distributions; integrate probability measures with normalized confusion matrices; rigorously define four critical domain types—easiest, hardest, dominant, and bottleneck—and prove that only specific scoring functions preserve weighted mean consistency. Our contributions include: (i) establishing a general theoretical foundation for multi-domain performance analysis; (ii) developing a visualization toolkit tailored to binary classification tasks; and (iii) enabling fine-grained, interpretable performance decomposition—thereby substantially enhancing transparency and practical utility in cross-domain evaluation. (149 words)

Analyzing multi-domain algorithm performance with user preference-based scoresDefining domain difficulty and importance based on user preferencesInvestigating weighted mean performance across different application domains

Existing software modeling datasets are often ad hoc constructions lacking rigorous quality assurance, leading to research findings that are difficult to reproduce, compare, and prone to bias. This work proposes the first benchmarking framework specifically designed for model-driven engineering, treating datasets themselves as first-class evaluation targets. By defining clear metrics for quality, representativeness, and task suitability, the framework establishes a unified platform that enables automated analysis of modeling datasets across multiple languages and formats. For the first time, this approach facilitates systematic evaluation of modeling datasets, substantially enhancing the reproducibility, fairness, and scientific rigor of research in the field.

benchmarkingdataset qualitymodel datasets

Diffusion models suffer from poor generalization in small-target-domain transfer learning, while test-time guidance methods incur high computational overhead and compromise sample diversity. To address these challenges, we propose DogFit, a domain-guided fine-tuning framework. Its core innovation lies in internalizing test-time guidance into the fine-tuning stage: lightweight conditional encoders dynamically inject domain-aware guidance offsets, and two scheduling strategies—late-start and truncation—implicitly balance fidelity and diversity during training. Built upon DiT/SiT architectures, DogFit leverages the strong marginal estimation capability of pretrained unconditional source-domain models and enables controllable generation with a single forward pass. Experiments across six target domains demonstrate that DogFit significantly outperforms existing guidance methods, achieving substantial improvements in FID and FDDINOV2 scores, reducing sampling TFLOPS by up to 2×, and incurring zero additional inference-time computation.

Balancing image fidelity and sample diversityEfficient transfer learning for diffusion modelsReducing computational cost in guided diffusion

Hot Scholars

RR

Raghavendra Ramachandra

Professor, Norwegian University of Science and Technology (NTNU), Norway
BiometricsImage/video analyticsDeep learningMachine Learning
CB

Christoph Busch

Professor for Biometrics, Norwegian University of Science and Technology (NTNU)
Biometrics
ZT

Zhaorui Tan

University of Liverpool, PHD student
GeneralizationText-to-ImageGenerative models
YH

Yahong Han

Professor of Computer Science, Tianjin University
Multimedia