align synthetic and real domains for zero-shot

Designs and implements data-generation, domain-alignment, and training methods so that models trained only on synthetic data can be applied to real-world data without further fine-tuning; this includes construction of synthetic-data pipelines, domain-randomization or alignment mechanisms, and training regimes that support zero-shot transfer. Develops evaluation protocols, metrics, and testing pipelines to measure zero-shot transfer performance and to analyze failure modes and the gap between synthetic and real domains.

alignsyntheticandreal

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.15
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Synthetic Dataset Evaluation Based on Generalized Cross Validation

Sep 14, 2025
ZS
Zhihang Song
🏛️ Tsinghua University

Existing synthetic data evaluation lacks unified, transferable quantitative metrics. This paper proposes a novel evaluation framework grounded in generalized cross-validation (GCV) and domain transfer learning. It constructs a cross-dataset performance matrix and defines two core metrics: *fidelity*, quantifying distributional similarity between synthetic and real data; and *generalization coverage*, measuring the task-transfer capability of synthetic data across diverse real-world source domains. The framework is model-agnostic and enables normalized, comparative evaluation of detectors such as YOLOv5s across heterogeneous datasets—including Virtual KITTI, KITTI, and BDD100K. Experiments demonstrate that the method effectively quantifies synthetic data quality, significantly enhancing evaluation generality, comparability, and utility for model optimization. It establishes a scalable, reproducible, and standardized evaluation paradigm for synthetic data development.

Evaluating synthetic dataset quality lacks standard frameworkProposing cross-validation and transfer learning for assessmentQuantifying simulation and transfer quality across domains

Beyond Real Data: Synthetic Data through the Lens of Regularization

Oct 09, 2025
AS
Amitis Shidani
🏛️ University of Oxford | Big Data Institute

In data-scarce real-world scenarios, synthetic data can enhance model generalization, yet excessive incorporation degrades performance due to distributional shift—e.g., increased Wasserstein distance—between synthetic and real domains. Method: We propose the first analytical framework grounded in algorithmic stability and regularization theory to quantify how the mixture ratio of synthetic to real data affects generalization error. Our analysis reveals, for the first time, a U-shaped relationship between test error and synthetic data proportion, and derives the theoretically optimal mixing ratio. Crucially, we incorporate the Wasserstein distance into the generalization bound for kernel ridge regression, extending it to domain adaptation settings. Results: Experiments on CIFAR-10 and clinical brain MRI datasets validate our theory: models trained with the predicted optimal ratio achieve significantly lower test error and demonstrate improved robustness and generalization—both in-domain and cross-domain.

Determining optimal synthetic-to-real data ratio for generalizationMitigating domain shift through strategic synthetic data blendingQuantifying the trade-off between synthetic and real data usage

SIDA: Synthetic Image Driven Zero-shot Domain Adaptation

Jul 24, 2025
YK
Ye-Chan Kim
🏛️ Hanyang University

Existing zero-shot domain adaptation (ZSDA) methods rely on textual descriptions to model target-domain style, which poorly captures complex real-world distribution shifts and incurs high alignment overhead and long adaptation latency. To address these limitations, we propose a synthetic-image-driven ZSDA framework: target-style synthetic images are generated via image translation—replacing hand-crafted text prompts—and serve as explicit style references. We introduce two novel modules—Domain Mix and Patch Style Transfer—to enable multi-style fusion and fine-grained local style transfer, respectively. Crucially, style features are extracted and transferred within the CLIP embedding space to preserve semantic consistency. Our approach significantly enhances modeling capability for severe domain shifts, achieving state-of-the-art performance across multiple ZSDA benchmarks—especially under challenging domain gaps—while reducing adaptation time. The method thus advances both efficiency and generalization in zero-shot domain adaptation.

Adapting models to target domains without target imagesOvercoming limitations of text-driven style feature simulationReducing adaptation time while capturing real-world variations

The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better

Jun 07, 2024
SG
Scott Geng
🏛️ University of Washington | Allen Institute for AI

Despite growing reliance on generative synthetic images (e.g., from Stable Diffusion) for data augmentation in image classification, their empirical effectiveness relative to real-world alternatives remains inadequately benchmarked. Method: This work systematically evaluates generative synthetic images against retrieval-based real images—obtained via CLIP cross-modal retrieval from LAION-2B—across multiple fine-grained classification tasks, using ViT and ResNet backbones for fine-tuning. Contribution/Results: Retrieval-based real images consistently match or significantly outperform synthetic counterparts across all tasks. Performance degradation in synthetic data is primarily attributed to generation artifacts and semantic misalignment. Crucially, this study establishes “simple retrieval” as a critical, empirically grounded baseline for evaluating synthetic data efficacy—challenging the prevailing overreliance on generative methods. To foster reproducibility and paradigmatic shift, the authors open-source all code, datasets, and models, advocating a transition in synthetic data research from “generation-first” to “utility-first” principles.

Generative ModelsImage RecognitionPerformance Comparison

Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World

Oct 22, 2024
JK
Joshua Kazdan
🏛️ Stanford University | Harvard

Do generative models inevitably suffer “model collapse” during large-scale pretraining with early-stage synthetic data? This paper systematically compares three synthetic-data training paradigms—replacement, accumulation, and constrained subset iteration—across Gaussian estimation, kernel density estimation, and language model fine-tuning. Methodologically, it introduces a generational iterative training framework, a multi-task benchmark suite, and dynamic test-loss modeling. Results demonstrate that “accumulation + full-dataset training” completely avoids collapse (test loss remains stable), whereas “constrained subset iteration” induces progressive performance degradation, and pure replacement inevitably collapses. These findings refute the monolithic assumption that synthetic data inherently causes collapse, instead establishing “data-evolution path dependence” as a new paradigm and empirically delineating safe operational boundaries for synthetic-data utilization.

Impact of synthetic data on model collapseLong-term stability of generative models with synthetic dataStrategies to manage synthetic data in training

Latest Papers

What's happening recently
View more

This study systematically evaluates the suitability and effectiveness of synthetic data across three canonical scenarios: data sharing, model training augmentation, and variance reduction in statistical estimation. By integrating formal modeling, theoretical analysis of generative models, and empirical case studies, the work presents the first comprehensive taxonomy of synthetic data applications and delineates their boundaries of applicability. The research elucidates both the potential and fundamental limitations of synthetic data in enhancing privacy preservation, model performance, and statistical stability. It further demonstrates that many existing or proposed use cases are misaligned with the intrinsic properties of synthetic data, thereby providing decision-makers with a principled theoretical framework to assess whether synthetic data is appropriate for addressing specific data availability challenges.

data augmentationdata sharingprivacy

The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.

AI trainingdata qualitydata volume

This study addresses the diminishing performance gains in large language model (LLM) training caused by redundancy and errors in synthetic data. It establishes, for the first time, a linear theoretical framework from a training dynamics perspective to characterize the value of synthetic data, explicitly defining optimal addition quantities and marginal utility. Based on this theory, we propose Training-Aware Target Coverage (TATC), a method that balances input coverage with error conditions to precisely select high-quality samples capable of effectively expanding directional coverage over target tasks for fine-tuning. Experiments demonstrate the validity of our theoretical analysis and show that TATC significantly outperforms existing baselines on the GSM8K mathematical reasoning task, substantially enhancing the performance of the Qwen2.5-Math model.

Data ValuationFine-tuningLarge Language Models

本文提出了一种针对数据可用性随时间演变的领域迁移学习问题(TrED),并探讨了现有方法在处理整个演化过程中的不足。

Data AvailabilityEvaluation CriterionEvolving Domains

Hot Scholars

AY

Alan Yuille

Professor of Cognitive Science and Computer Science, Johns Hopkins University
Computer VisionComputational Models of Mind and BrainMachine Learning
HJ

Hong Jia

Lecturer (Assistant Professor), University of Auckland; University of Melbourne
On-Device MLHuman-Centred AIMobile ComputingMobile Health
YP

Yuxin Peng

Peking University
Cross-media Analysis and ReasoningImage & Video Understanding and RetrievalMachine Learning and Artificial Intelligence
YL

Yang Liu

Peking University
Computer VisionMulti-modal Learning