Score
Designs and implements pretraining regimes for machine learning models that rely on synthetic data produced by simulation; this includes building data-generating procedures, creating large-scale simulated datasets, and configuring pretraining objectives and efficient training pipelines. Analyzes how pretrained models transfer to real data, measures sensitivity to simulation assumptions, and optimizes compute and data-generation choices to reduce dependence on scarce or revised real-world datasets.
This study addresses the lack of systematic synthesis at the intersection of artificial intelligence (AI) and modeling and simulation (M&S) by proposing, for the first time, a structured framework based on the full M&S lifecycle—encompassing model construction, input modeling, execution, experimentation, validation, and output analysis. It elucidates the bidirectional integration mechanisms between AI and simulation: how AI enhances or substitutes traditional simulation components, and how simulation supports AI training and evaluation. Incorporating generative AI technologies such as large language models, the paper identifies representative application paradigms and integration approaches across each phase, synthesizes key achievements, and presents a conceptual roadmap tailored to the rapidly evolving ecosystem, while highlighting current limitations and open research challenges.
Subsymbolic AI struggles with training under few-shot and low-quality data, while existing virtual simulation approaches lack systematic, standards-aligned frameworks. Method: We conduct a systematic literature review covering 22 state-of-the-art works and propose, for the first time, a unified reference framework for digital twin–driven AI simulation—deeply integrating digital twins with AI agents to establish a closed-loop, cyber-physical data orchestration mechanism. We further achieve systematic alignment of this framework with the ISO 23247 international standard for digital twins. Contribution/Results: We distill key technological evolution trends, identify five core challenges and several open research directions, and deliver a reusable architectural guideline and reference framework. This work provides a standardized, methodology-driven foundation for high-fidelity AI simulation, advancing both theoretical rigor and practical deployability in industrial AI applications.
Do generative models inevitably suffer “model collapse” during large-scale pretraining with early-stage synthetic data? This paper systematically compares three synthetic-data training paradigms—replacement, accumulation, and constrained subset iteration—across Gaussian estimation, kernel density estimation, and language model fine-tuning. Methodologically, it introduces a generational iterative training framework, a multi-task benchmark suite, and dynamic test-loss modeling. Results demonstrate that “accumulation + full-dataset training” completely avoids collapse (test loss remains stable), whereas “constrained subset iteration” induces progressive performance degradation, and pure replacement inevitably collapses. These findings refute the monolithic assumption that synthetic data inherently causes collapse, instead establishing “data-evolution path dependence” as a new paradigm and empirically delineating safe operational boundaries for synthetic-data utilization.
The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.
This study systematically evaluates the suitability and effectiveness of synthetic data across three canonical scenarios: data sharing, model training augmentation, and variance reduction in statistical estimation. By integrating formal modeling, theoretical analysis of generative models, and empirical case studies, the work presents the first comprehensive taxonomy of synthetic data applications and delineates their boundaries of applicability. The research elucidates both the potential and fundamental limitations of synthetic data in enhancing privacy preservation, model performance, and statistical stability. It further demonstrates that many existing or proposed use cases are misaligned with the intrinsic properties of synthetic data, thereby providing decision-makers with a principled theoretical framework to assess whether synthetic data is appropriate for addressing specific data availability challenges.
In the era of large language models, selecting and fine-tuning pre-trained models remains challenging due to the absence of efficient adaptation mechanisms across vast model zoos and severe limitations in labeled data, hindering few-shot generalization. Method: (1) We systematically integrate meta-learning into the deep learning pipeline, constructing a task-prior-driven pipeline ranking surrogate model; (2) we quantitatively characterize the critical role of data augmentation in self-supervised learning; (3) we propose a differentiable neural synthetic data generator, replacing conventional reinforcement learning–based approaches. Contribution/Results: Our framework significantly outperforms human-crafted fine-tuning baselines on standard CV and NLP benchmarks. It achieves substantial gains in few-shot settings and enables zero-shot cross-environment generalization of the synthetic data generator, demonstrating robust adaptability without domain-specific retraining.
Machine learning models often suffer from poor real-world data quality, limited sample availability, and stringent privacy regulations—leading to underfitting and hindered data sharing. This paper presents a systematic survey of generative synthetic data techniques across five domains: computer vision, speech, natural language processing, healthcare, and business analytics. We unify the modeling of generation mechanisms, privacy preservation, and fairness considerations. Methodologically, we propose a novel cross-modal evaluation framework that integrates GANs, VAEs, diffusion models, autoregressive models, and differential privacy–enhanced approaches—thereby establishing theoretical bounds and identifying practical gaps in the privacy–utility trade-off. We further introduce a unified taxonomy of synthetic data and identify six fundamental challenges and four emerging opportunities. This work delivers the first multidisciplinary academic roadmap toward trustworthy, standardized synthetic data practice.
This work addresses the risks of bias in statistical inference when using synthetic data generated by modern generative AI models—such as diffusion models, GANs, and large language models—due to model misspecification, underestimation of uncertainty, and insufficient generalization. It presents the first systematic integration of generative modeling with statistical inference theory, clarifying the assumptions and conditions under which synthetic data can reliably support scientific discovery. By unifying uncertainty quantification, model diagnostics, and downstream task analysis, the study proposes a principled framework that delineates effective usage guidelines, identifies critical failure modes, and offers practical recommendations for researchers and developers, along with directions for future research.
This study investigates the fundamental limitations of synthetic data augmentation in enhancing sample information for statistical inference, with particular emphasis on its theoretical constraints when incorporating prior knowledge. Treating synthetic data as a model of the prior, we formally define the synthetic distribution within both maximum likelihood and Bayesian frameworks. By integrating Fisher information, information theory, and statistical decision theory, we establish an intrinsic upper bound on the marginal information gain achievable through synthetic data. Our analysis reveals that naive prior specifications lack epistemic justification and generally fail to improve conventional inferential performance. Nevertheless, under the training/test partition paradigm, synthetic data can effectively regularize high-dimensional model spaces by imposing structural constraints, thereby serving a beneficial regularization role despite its limited informational contribution.
This study addresses the diminishing performance gains in large language model (LLM) training caused by redundancy and errors in synthetic data. It establishes, for the first time, a linear theoretical framework from a training dynamics perspective to characterize the value of synthetic data, explicitly defining optimal addition quantities and marginal utility. Based on this theory, we propose Training-Aware Target Coverage (TATC), a method that balances input coverage with error conditions to precisely select high-quality samples capable of effectively expanding directional coverage over target tasks for fine-tuning. Experiments demonstrate the validity of our theoretical analysis and show that TATC significantly outperforms existing baselines on the GSM8K mathematical reasoning task, substantially enhancing the performance of the Qwen2.5-Math model.
This work addresses the challenges of real-world data scarcity, high acquisition costs, and privacy sensitivity in multimodal AI training by introducing Simula, a novel framework that pioneers inference-driven synthetic data generation without requiring any seed data. By integrating an agent-based architecture with a controllable generation pipeline, Simula enables fine-grained control over data characteristics and computational resource allocation, substantially enhancing the interpretability and scalability of synthetic data. Through a comprehensive multidimensional evaluation protocol, the framework simultaneously validates both the intrinsic quality of the generated data and its effectiveness in downstream tasks across multiple benchmarks, offering a practical pathway and design paradigm for AI development under data-constrained conditions.
It remains unclear whether synthetic data generated by generative models can improve classifier generalization, and existing heuristic selection criteria lack theoretical foundations. Method: We systematically analyze the impact of synthetic data selection on generalization error from a high-dimensional regression perspective, identifying covariance shift—not mean shift—as the key limiting factor. Based on this insight, we propose the “covariance matching” principle and prove its optimality under mild conditions. The framework is applicable to deep neural networks and mainstream generative models. Contribution/Results: Through rigorous theoretical analysis, linear model studies, and extensive empirical validation across diverse architectures, datasets, and generative models (e.g., GANs, VAEs, diffusion models), we demonstrate that covariance matching consistently outperforms existing synthetic data selection strategies, yielding stable and significant improvements in classifier prediction performance.