Score
Designs and implements data-generation, domain-alignment, and training methods so that models trained only on synthetic data can be applied to real-world data without further fine-tuning; this includes construction of synthetic-data pipelines, domain-randomization or alignment mechanisms, and training regimes that support zero-shot transfer. Develops evaluation protocols, metrics, and testing pipelines to measure zero-shot transfer performance and to analyze failure modes and the gap between synthetic and real domains.
This paper addresses the research gap in applying Contrastive Language–Image Pretraining (CLIP) to domain generalization (DG) and domain adaptation (DA). Methodologically, it establishes a unified taxonomy by systematically analyzing prevailing paradigms—including prompt optimization, backbone feature reuse, and source-available/source-free transfer—thereby elucidating CLIP’s zero-shot cross-domain transfer mechanisms and pathways to enhanced robustness. The analysis identifies three critical bottlenecks: overfitting, insufficient domain diversity, and computational inefficiency. To overcome these, the work integrates key techniques such as prompt learning, feature alignment, knowledge distillation, and domain-invariant representation learning. The contributions include both a rigorous methodological framework for CLIP-based DG/DA and actionable insights into architectural design and training strategies. Collectively, this study provides theoretical foundations and practical guidelines for developing more generalizable and deployable cross-domain vision models.
Existing synthetic data evaluation lacks unified, transferable quantitative metrics. This paper proposes a novel evaluation framework grounded in generalized cross-validation (GCV) and domain transfer learning. It constructs a cross-dataset performance matrix and defines two core metrics: *fidelity*, quantifying distributional similarity between synthetic and real data; and *generalization coverage*, measuring the task-transfer capability of synthetic data across diverse real-world source domains. The framework is model-agnostic and enables normalized, comparative evaluation of detectors such as YOLOv5s across heterogeneous datasets—including Virtual KITTI, KITTI, and BDD100K. Experiments demonstrate that the method effectively quantifies synthetic data quality, significantly enhancing evaluation generality, comparability, and utility for model optimization. It establishes a scalable, reproducible, and standardized evaluation paradigm for synthetic data development.
In data-scarce real-world scenarios, synthetic data can enhance model generalization, yet excessive incorporation degrades performance due to distributional shift—e.g., increased Wasserstein distance—between synthetic and real domains. Method: We propose the first analytical framework grounded in algorithmic stability and regularization theory to quantify how the mixture ratio of synthetic to real data affects generalization error. Our analysis reveals, for the first time, a U-shaped relationship between test error and synthetic data proportion, and derives the theoretically optimal mixing ratio. Crucially, we incorporate the Wasserstein distance into the generalization bound for kernel ridge regression, extending it to domain adaptation settings. Results: Experiments on CIFAR-10 and clinical brain MRI datasets validate our theory: models trained with the predicted optimal ratio achieve significantly lower test error and demonstrate improved robustness and generalization—both in-domain and cross-domain.
Existing zero-shot domain adaptation (ZSDA) methods rely on textual descriptions to model target-domain style, which poorly captures complex real-world distribution shifts and incurs high alignment overhead and long adaptation latency. To address these limitations, we propose a synthetic-image-driven ZSDA framework: target-style synthetic images are generated via image translation—replacing hand-crafted text prompts—and serve as explicit style references. We introduce two novel modules—Domain Mix and Patch Style Transfer—to enable multi-style fusion and fine-grained local style transfer, respectively. Crucially, style features are extracted and transferred within the CLIP embedding space to preserve semantic consistency. Our approach significantly enhances modeling capability for severe domain shifts, achieving state-of-the-art performance across multiple ZSDA benchmarks—especially under challenging domain gaps—while reducing adaptation time. The method thus advances both efficiency and generalization in zero-shot domain adaptation.
Despite growing reliance on generative synthetic images (e.g., from Stable Diffusion) for data augmentation in image classification, their empirical effectiveness relative to real-world alternatives remains inadequately benchmarked. Method: This work systematically evaluates generative synthetic images against retrieval-based real images—obtained via CLIP cross-modal retrieval from LAION-2B—across multiple fine-grained classification tasks, using ViT and ResNet backbones for fine-tuning. Contribution/Results: Retrieval-based real images consistently match or significantly outperform synthetic counterparts across all tasks. Performance degradation in synthetic data is primarily attributed to generation artifacts and semantic misalignment. Crucially, this study establishes “simple retrieval” as a critical, empirically grounded baseline for evaluating synthetic data efficacy—challenging the prevailing overreliance on generative methods. To foster reproducibility and paradigmatic shift, the authors open-source all code, datasets, and models, advocating a transition in synthetic data research from “generation-first” to “utility-first” principles.
Do generative models inevitably suffer “model collapse” during large-scale pretraining with early-stage synthetic data? This paper systematically compares three synthetic-data training paradigms—replacement, accumulation, and constrained subset iteration—across Gaussian estimation, kernel density estimation, and language model fine-tuning. Methodologically, it introduces a generational iterative training framework, a multi-task benchmark suite, and dynamic test-loss modeling. Results demonstrate that “accumulation + full-dataset training” completely avoids collapse (test loss remains stable), whereas “constrained subset iteration” induces progressive performance degradation, and pure replacement inevitably collapses. These findings refute the monolithic assumption that synthetic data inherently causes collapse, instead establishing “data-evolution path dependence” as a new paradigm and empirically delineating safe operational boundaries for synthetic-data utilization.
This study systematically evaluates the suitability and effectiveness of synthetic data across three canonical scenarios: data sharing, model training augmentation, and variance reduction in statistical estimation. By integrating formal modeling, theoretical analysis of generative models, and empirical case studies, the work presents the first comprehensive taxonomy of synthetic data applications and delineates their boundaries of applicability. The research elucidates both the potential and fundamental limitations of synthetic data in enhancing privacy preservation, model performance, and statistical stability. It further demonstrates that many existing or proposed use cases are misaligned with the intrinsic properties of synthetic data, thereby providing decision-makers with a principled theoretical framework to assess whether synthetic data is appropriate for addressing specific data availability challenges.
The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.
本文解决如何用最少的真实数据测试合成数据集对AI模型训练效果的影响,提出了一种自适应e-过程符号翻转测试方法。
This study addresses the diminishing performance gains in large language model (LLM) training caused by redundancy and errors in synthetic data. It establishes, for the first time, a linear theoretical framework from a training dynamics perspective to characterize the value of synthetic data, explicitly defining optimal addition quantities and marginal utility. Based on this theory, we propose Training-Aware Target Coverage (TATC), a method that balances input coverage with error conditions to precisely select high-quality samples capable of effectively expanding directional coverage over target tasks for fine-tuning. Experiments demonstrate the validity of our theoretical analysis and show that TATC significantly outperforms existing baselines on the GSM8K mathematical reasoning task, substantially enhancing the performance of the Qwen2.5-Math model.
本文提出了一种针对数据可用性随时间演变的领域迁移学习问题(TrED),并探讨了现有方法在处理整个演化过程中的不足。