align synthetic and real domains for zero-shot

Designs and implements data-generation, domain-alignment, and training methods so that models trained only on synthetic data can be applied to real-world data without further fine-tuning; this includes construction of synthetic-data pipelines, domain-randomization or alignment mechanisms, and training regimes that support zero-shot transfer. Develops evaluation protocols, metrics, and testing pipelines to measure zero-shot transfer performance and to analyze failure modes and the gap between synthetic and real domains.

alignsyntheticandreal

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.11
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Synthetic Dataset Evaluation Based on Generalized Cross Validation

Sep 14, 2025
ZS
Zhihang Song
🏛️ Tsinghua University

Existing synthetic data evaluation lacks unified, transferable quantitative metrics. This paper proposes a novel evaluation framework grounded in generalized cross-validation (GCV) and domain transfer learning. It constructs a cross-dataset performance matrix and defines two core metrics: *fidelity*, quantifying distributional similarity between synthetic and real data; and *generalization coverage*, measuring the task-transfer capability of synthetic data across diverse real-world source domains. The framework is model-agnostic and enables normalized, comparative evaluation of detectors such as YOLOv5s across heterogeneous datasets—including Virtual KITTI, KITTI, and BDD100K. Experiments demonstrate that the method effectively quantifies synthetic data quality, significantly enhancing evaluation generality, comparability, and utility for model optimization. It establishes a scalable, reproducible, and standardized evaluation paradigm for synthetic data development.

Evaluating synthetic dataset quality lacks standard frameworkProposing cross-validation and transfer learning for assessmentQuantifying simulation and transfer quality across domains

Beyond Real Data: Synthetic Data through the Lens of Regularization

Oct 09, 2025
AS
Amitis Shidani
🏛️ University of Oxford | Big Data Institute

In data-scarce real-world scenarios, synthetic data can enhance model generalization, yet excessive incorporation degrades performance due to distributional shift—e.g., increased Wasserstein distance—between synthetic and real domains. Method: We propose the first analytical framework grounded in algorithmic stability and regularization theory to quantify how the mixture ratio of synthetic to real data affects generalization error. Our analysis reveals, for the first time, a U-shaped relationship between test error and synthetic data proportion, and derives the theoretically optimal mixing ratio. Crucially, we incorporate the Wasserstein distance into the generalization bound for kernel ridge regression, extending it to domain adaptation settings. Results: Experiments on CIFAR-10 and clinical brain MRI datasets validate our theory: models trained with the predicted optimal ratio achieve significantly lower test error and demonstrate improved robustness and generalization—both in-domain and cross-domain.

Determining optimal synthetic-to-real data ratio for generalizationMitigating domain shift through strategic synthetic data blendingQuantifying the trade-off between synthetic and real data usage

SIDA: Synthetic Image Driven Zero-shot Domain Adaptation

Jul 24, 2025
YK
Ye-Chan Kim
🏛️ Hanyang University

Existing zero-shot domain adaptation (ZSDA) methods rely on textual descriptions to model target-domain style, which poorly captures complex real-world distribution shifts and incurs high alignment overhead and long adaptation latency. To address these limitations, we propose a synthetic-image-driven ZSDA framework: target-style synthetic images are generated via image translation—replacing hand-crafted text prompts—and serve as explicit style references. We introduce two novel modules—Domain Mix and Patch Style Transfer—to enable multi-style fusion and fine-grained local style transfer, respectively. Crucially, style features are extracted and transferred within the CLIP embedding space to preserve semantic consistency. Our approach significantly enhances modeling capability for severe domain shifts, achieving state-of-the-art performance across multiple ZSDA benchmarks—especially under challenging domain gaps—while reducing adaptation time. The method thus advances both efficiency and generalization in zero-shot domain adaptation.

Adapting models to target domains without target imagesOvercoming limitations of text-driven style feature simulationReducing adaptation time while capturing real-world variations

The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better

Jun 07, 2024
SG
Scott Geng
🏛️ University of Washington | Allen Institute for AI

Despite growing reliance on generative synthetic images (e.g., from Stable Diffusion) for data augmentation in image classification, their empirical effectiveness relative to real-world alternatives remains inadequately benchmarked. Method: This work systematically evaluates generative synthetic images against retrieval-based real images—obtained via CLIP cross-modal retrieval from LAION-2B—across multiple fine-grained classification tasks, using ViT and ResNet backbones for fine-tuning. Contribution/Results: Retrieval-based real images consistently match or significantly outperform synthetic counterparts across all tasks. Performance degradation in synthetic data is primarily attributed to generation artifacts and semantic misalignment. Crucially, this study establishes “simple retrieval” as a critical, empirically grounded baseline for evaluating synthetic data efficacy—challenging the prevailing overreliance on generative methods. To foster reproducibility and paradigmatic shift, the authors open-source all code, datasets, and models, advocating a transition in synthetic data research from “generation-first” to “utility-first” principles.

Generative ModelsImage RecognitionPerformance Comparison

Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World

Oct 22, 2024
JK
Joshua Kazdan
🏛️ Stanford University | Harvard

Do generative models inevitably suffer “model collapse” during large-scale pretraining with early-stage synthetic data? This paper systematically compares three synthetic-data training paradigms—replacement, accumulation, and constrained subset iteration—across Gaussian estimation, kernel density estimation, and language model fine-tuning. Methodologically, it introduces a generational iterative training framework, a multi-task benchmark suite, and dynamic test-loss modeling. Results demonstrate that “accumulation + full-dataset training” completely avoids collapse (test loss remains stable), whereas “constrained subset iteration” induces progressive performance degradation, and pure replacement inevitably collapses. These findings refute the monolithic assumption that synthetic data inherently causes collapse, instead establishing “data-evolution path dependence” as a new paradigm and empirically delineating safe operational boundaries for synthetic-data utilization.

Impact of synthetic data on model collapseLong-term stability of generative models with synthetic dataStrategies to manage synthetic data in training

Latest Papers

What's happening recently
View more

This study systematically evaluates the suitability and effectiveness of synthetic data across three canonical scenarios: data sharing, model training augmentation, and variance reduction in statistical estimation. By integrating formal modeling, theoretical analysis of generative models, and empirical case studies, the work presents the first comprehensive taxonomy of synthetic data applications and delineates their boundaries of applicability. The research elucidates both the potential and fundamental limitations of synthetic data in enhancing privacy preservation, model performance, and statistical stability. It further demonstrates that many existing or proposed use cases are misaligned with the intrinsic properties of synthetic data, thereby providing decision-makers with a principled theoretical framework to assess whether synthetic data is appropriate for addressing specific data availability challenges.

data augmentationdata sharingprivacy

Using Synthetic Data to estimate the True Error is theoretically and practically doable

Nov 02, 2025
HH
Hai Hoang Thanh
🏛️ Hanoi University of Science and Technology | Vietnam Posts and Telecommunications Group

Accurately estimating model test error under scarce labeled data remains challenging. Method: This paper proposes a novel error estimation paradigm leveraging high-quality synthetic data. Theoretically, we derive a new generalization error upper bound incorporating generator quality constraints, quantifying for the first time how generative model fidelity critically affects estimation bias. Methodologically, we design an interpretable and optimization-friendly synthetic sample construction strategy that jointly leverages generative modeling and generalization theory to enhance assessment reliability. Results: Extensive experiments on both synthetic and real-world tabular datasets demonstrate that our approach consistently outperforms existing baselines, achieving significant and robust improvements in both accuracy and stability of error estimation.

Developing generalization bounds incorporating synthetic data for error estimation.Estimating model test error using synthetic data under limited labeled samples.Optimizing synthetic data generation to improve evaluation accuracy and reliability.

The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.

AI trainingdata qualitydata volume

Hot Scholars

AY

Alan Yuille

Professor of Cognitive Science and Computer Science, Johns Hopkins University
Computer VisionComputational Models of Mind and BrainMachine Learning
YP

Yuxin Peng

Peking University
Cross-media Analysis and ReasoningImage & Video Understanding and RetrievalMachine Learning and Artificial Intelligence
YL

Yang Liu

Peking University
Computer VisionMulti-modal Learning
YY

Yun Ye

Intel
Computer VisionDeep LearningSemiconductor Physics