Score
Designs and implements data-generation, domain-alignment, and training methods so that models trained only on synthetic data can be applied to real-world data without further fine-tuning; this includes construction of synthetic-data pipelines, domain-randomization or alignment mechanisms, and training regimes that support zero-shot transfer. Develops evaluation protocols, metrics, and testing pipelines to measure zero-shot transfer performance and to analyze failure modes and the gap between synthetic and real domains.
This paper addresses the research gap in applying Contrastive Language–Image Pretraining (CLIP) to domain generalization (DG) and domain adaptation (DA). Methodologically, it establishes a unified taxonomy by systematically analyzing prevailing paradigms—including prompt optimization, backbone feature reuse, and source-available/source-free transfer—thereby elucidating CLIP’s zero-shot cross-domain transfer mechanisms and pathways to enhanced robustness. The analysis identifies three critical bottlenecks: overfitting, insufficient domain diversity, and computational inefficiency. To overcome these, the work integrates key techniques such as prompt learning, feature alignment, knowledge distillation, and domain-invariant representation learning. The contributions include both a rigorous methodological framework for CLIP-based DG/DA and actionable insights into architectural design and training strategies. Collectively, this study provides theoretical foundations and practical guidelines for developing more generalizable and deployable cross-domain vision models.
Existing synthetic data evaluation lacks unified, transferable quantitative metrics. This paper proposes a novel evaluation framework grounded in generalized cross-validation (GCV) and domain transfer learning. It constructs a cross-dataset performance matrix and defines two core metrics: *fidelity*, quantifying distributional similarity between synthetic and real data; and *generalization coverage*, measuring the task-transfer capability of synthetic data across diverse real-world source domains. The framework is model-agnostic and enables normalized, comparative evaluation of detectors such as YOLOv5s across heterogeneous datasets—including Virtual KITTI, KITTI, and BDD100K. Experiments demonstrate that the method effectively quantifies synthetic data quality, significantly enhancing evaluation generality, comparability, and utility for model optimization. It establishes a scalable, reproducible, and standardized evaluation paradigm for synthetic data development.
In data-scarce real-world scenarios, synthetic data can enhance model generalization, yet excessive incorporation degrades performance due to distributional shift—e.g., increased Wasserstein distance—between synthetic and real domains. Method: We propose the first analytical framework grounded in algorithmic stability and regularization theory to quantify how the mixture ratio of synthetic to real data affects generalization error. Our analysis reveals, for the first time, a U-shaped relationship between test error and synthetic data proportion, and derives the theoretically optimal mixing ratio. Crucially, we incorporate the Wasserstein distance into the generalization bound for kernel ridge regression, extending it to domain adaptation settings. Results: Experiments on CIFAR-10 and clinical brain MRI datasets validate our theory: models trained with the predicted optimal ratio achieve significantly lower test error and demonstrate improved robustness and generalization—both in-domain and cross-domain.
Existing zero-shot domain adaptation (ZSDA) methods rely on textual descriptions to model target-domain style, which poorly captures complex real-world distribution shifts and incurs high alignment overhead and long adaptation latency. To address these limitations, we propose a synthetic-image-driven ZSDA framework: target-style synthetic images are generated via image translation—replacing hand-crafted text prompts—and serve as explicit style references. We introduce two novel modules—Domain Mix and Patch Style Transfer—to enable multi-style fusion and fine-grained local style transfer, respectively. Crucially, style features are extracted and transferred within the CLIP embedding space to preserve semantic consistency. Our approach significantly enhances modeling capability for severe domain shifts, achieving state-of-the-art performance across multiple ZSDA benchmarks—especially under challenging domain gaps—while reducing adaptation time. The method thus advances both efficiency and generalization in zero-shot domain adaptation.
Despite growing reliance on generative synthetic images (e.g., from Stable Diffusion) for data augmentation in image classification, their empirical effectiveness relative to real-world alternatives remains inadequately benchmarked. Method: This work systematically evaluates generative synthetic images against retrieval-based real images—obtained via CLIP cross-modal retrieval from LAION-2B—across multiple fine-grained classification tasks, using ViT and ResNet backbones for fine-tuning. Contribution/Results: Retrieval-based real images consistently match or significantly outperform synthetic counterparts across all tasks. Performance degradation in synthetic data is primarily attributed to generation artifacts and semantic misalignment. Crucially, this study establishes “simple retrieval” as a critical, empirically grounded baseline for evaluating synthetic data efficacy—challenging the prevailing overreliance on generative methods. To foster reproducibility and paradigmatic shift, the authors open-source all code, datasets, and models, advocating a transition in synthetic data research from “generation-first” to “utility-first” principles.
Do generative models inevitably suffer “model collapse” during large-scale pretraining with early-stage synthetic data? This paper systematically compares three synthetic-data training paradigms—replacement, accumulation, and constrained subset iteration—across Gaussian estimation, kernel density estimation, and language model fine-tuning. Methodologically, it introduces a generational iterative training framework, a multi-task benchmark suite, and dynamic test-loss modeling. Results demonstrate that “accumulation + full-dataset training” completely avoids collapse (test loss remains stable), whereas “constrained subset iteration” induces progressive performance degradation, and pure replacement inevitably collapses. These findings refute the monolithic assumption that synthetic data inherently causes collapse, instead establishing “data-evolution path dependence” as a new paradigm and empirically delineating safe operational boundaries for synthetic-data utilization.
This study systematically evaluates the suitability and effectiveness of synthetic data across three canonical scenarios: data sharing, model training augmentation, and variance reduction in statistical estimation. By integrating formal modeling, theoretical analysis of generative models, and empirical case studies, the work presents the first comprehensive taxonomy of synthetic data applications and delineates their boundaries of applicability. The research elucidates both the potential and fundamental limitations of synthetic data in enhancing privacy preservation, model performance, and statistical stability. It further demonstrates that many existing or proposed use cases are misaligned with the intrinsic properties of synthetic data, thereby providing decision-makers with a principled theoretical framework to assess whether synthetic data is appropriate for addressing specific data availability challenges.
Accurately estimating model test error under scarce labeled data remains challenging. Method: This paper proposes a novel error estimation paradigm leveraging high-quality synthetic data. Theoretically, we derive a new generalization error upper bound incorporating generator quality constraints, quantifying for the first time how generative model fidelity critically affects estimation bias. Methodologically, we design an interpretable and optimization-friendly synthetic sample construction strategy that jointly leverages generative modeling and generalization theory to enhance assessment reliability. Results: Extensive experiments on both synthetic and real-world tabular datasets demonstrate that our approach consistently outperforms existing baselines, achieving significant and robust improvements in both accuracy and stability of error estimation.
The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.