Score
Designs and builds synthetic datasets and end-to-end generation pipelines that produce realistic, diverse, hierarchical, and scalable artificial instances for training and evaluation, using methods such as domain-randomized rendering, non‑parametric simulation, hierarchical/random-hierarchy models, generative dataset expansion, and LLM-assisted synthesis. Also curates and analyzes the statistical fidelity, structural diversity, and scalability of those datasets (instance-level generation, dataset construction, and dataset curation) to ensure they preserve relevant real-world nuance and desired variation.
Machine learning models often suffer from poor real-world data quality, limited sample availability, and stringent privacy regulations—leading to underfitting and hindered data sharing. This paper presents a systematic survey of generative synthetic data techniques across five domains: computer vision, speech, natural language processing, healthcare, and business analytics. We unify the modeling of generation mechanisms, privacy preservation, and fairness considerations. Methodologically, we propose a novel cross-modal evaluation framework that integrates GANs, VAEs, diffusion models, autoregressive models, and differential privacy–enhanced approaches—thereby establishing theoretical bounds and identifying practical gaps in the privacy–utility trade-off. We further introduce a unified taxonomy of synthetic data and identify six fundamental challenges and four emerging opportunities. This work delivers the first multidisciplinary academic roadmap toward trustworthy, standardized synthetic data practice.
This work addresses the challenges of real-world data scarcity, high acquisition costs, and privacy sensitivity in multimodal AI training by introducing Simula, a novel framework that pioneers inference-driven synthetic data generation without requiring any seed data. By integrating an agent-based architecture with a controllable generation pipeline, Simula enables fine-grained control over data characteristics and computational resource allocation, substantially enhancing the interpretability and scalability of synthetic data. Through a comprehensive multidimensional evaluation protocol, the framework simultaneously validates both the intrinsic quality of the generated data and its effectiveness in downstream tasks across multiple benchmarks, offering a practical pathway and design paradigm for AI development under data-constrained conditions.
The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.
To address core challenges in data mining—including data scarcity, privacy sensitivity, and high annotation costs—this paper proposes a task-oriented synthetic data generation paradigm. Methodologically, it systematically integrates state-of-the-art generative models—large language models, diffusion models, and generative adversarial networks—within a unified evaluation framework and reusable practical guidelines. Key contributions include: (i) the first coordinated application of multimodal generative models to data mining tasks, jointly optimizing data fidelity, statistical utility, and privacy preservation; (ii) an end-to-end synthetic data quality assessment metric suite; and (iii) open-sourced tutorials and a dedicated tool website. Experiments demonstrate that the generated synthetic data significantly improves downstream model performance (average +12.3% F1 score) while satisfying rigorous privacy constraints such as differential privacy. This work provides both a methodological foundation and an engineering blueprint for trustworthy, scalable, data-driven research in the GenAI era.
Instance segmentation of unseen objects in cluttered desktop scenes requires both modal (visible) and amodal (full-object) masks, yet acquiring large-scale, diverse, and accurately annotated real-world data remains challenging. Method: We propose the first end-to-end synthetic data generation framework tailored for amodal segmentation in desktop scenarios. Built upon NVIDIA Isaac Sim Replicator Composer, it implements a Python pipeline that automatically renders photorealistic 3D desktop scenes with varied materials, lighting, and textures, while simultaneously generating rich annotations—including semantic/instance masks, depth maps, occlusion masks, and amodal masks. The framework supports user-defined annotation types and eliminates manual labeling entirely. Results: Evaluated on the OSD-Amodal dataset, UOAIS-Net trained exclusively on our synthetic data achieves state-of-the-art sim-to-real transfer performance. We publicly release the code, a representative synthetic dataset, and demonstration videos.
To address the pervasive diminishing returns in synthetic data augmentation, this paper proposes DP, a dynamic synthetic data generation framework inspired by the pedagogical principle of “deliberate practice.” Departing from conventional “generate-then-prune” paradigms, DP directly targets high-value sample distributions via three core mechanisms: dynamic difficulty adjustment, information-theoretic sample generation, and theory-guided sample selection. It is the first work to formally incorporate human learning principles—specifically, challenge-based training—into synthetic data generation and theoretically proves its positive impact on model scaling laws. Implemented as a lightweight plugin, DP seamlessly integrates with mainstream diffusion models and large language models. Experiments demonstrate significant efficiency gains: on ImageNet-100, it reduces required sample count and training iterations by 3.4× and 6×, respectively; on ImageNet-1K, reductions reach 8× in samples and 30% in iterations—while consistently outperforming state-of-the-art methods across all benchmarks.
This work addresses the risks of bias in statistical inference when using synthetic data generated by modern generative AI models—such as diffusion models, GANs, and large language models—due to model misspecification, underestimation of uncertainty, and insufficient generalization. It presents the first systematic integration of generative modeling with statistical inference theory, clarifying the assumptions and conditions under which synthetic data can reliably support scientific discovery. By unifying uncertainty quantification, model diagnostics, and downstream task analysis, the study proposes a principled framework that delineates effective usage guidelines, identifies critical failure modes, and offers practical recommendations for researchers and developers, along with directions for future research.
This study addresses the limitations of traditional Monte Carlo simulations, which rely on ad hoc assumptions and struggle to generate data reflecting realistic multilevel structures, thereby compromising the validity of quantitative method evaluations. To overcome this, the authors propose the first six-stage workflow integrating generative AI with multilevel data simulation. They innovatively adapt diffusion models and generative adversarial networks (GANs) to accommodate hierarchical data structures and introduce a comprehensive synthetic data quality assessment framework that ensures both within-table and cross-table consistency. Empirical experiments on real-world social science datasets demonstrate that the proposed approach substantially enhances the realism and reliability of Monte Carlo simulations, outperforming conventional strategies and providing a more empirically grounded benchmark for evaluating predictive performance and parameter recovery in quantitative methods.
This work proposes a GAN-inspired privacy-preserving synthetic data generation method that avoids direct access to original data during training. Instead, it leverages fuzz testing to produce candidate samples and iteratively refines them through a discriminator-guided feedback loop combined with statistical distribution constraints to approximate the original data distribution. By innovatively integrating fuzz testing, adversarial discrimination, and indirect constraint mechanisms, the approach achieves strong privacy guarantees—effectively resisting membership inference and data reconstruction attacks—while preserving high data utility. Extensive experiments on four benchmark datasets demonstrate that the proposed method strikes a superior balance between privacy protection and data fidelity compared to existing techniques.
This work addresses a critical yet overlooked limitation of modern generative models: their tendency to produce class-typical samples at the expense of intra-class diversity, thereby diminishing the utility of synthetic data in downstream tasks. The study is the first to formally characterize this structural bias and introduces a novel post-hoc filtering mechanism that requires neither retraining nor generator-specific modifications. By partitioning real classes into homogeneous typical (HO) and heterogeneous non-redundant (HE) subsets, the method selects high-quality synthetic samples through a fidelity–diversity criterion that combines semantic alignment scores with redundancy penalties. Evaluated across multiple benchmarks, the approach consistently outperforms existing data selection strategies—achieving performance on par with real data using only 60% of the synthetic samples—and provides consistent gains even when applied to strong generative models in both classification and segmentation tasks.
This work addresses critical limitations in existing large language model (LLM)-based 3D scene generation methods for agricultural applications, including insufficient domain-specific knowledge, lack of validation mechanisms, and inadequate modularity, which collectively constrain controllability and scalability. To overcome these challenges, we propose a modular multi-LLM pipeline that integrates agricultural domain knowledge, few-shot prompting, retrieval-augmented generation (RAG), and Unreal Engine APIs to automatically construct realistic agricultural simulation environments. The architecture enables intermediate validation, structured data handling, and flexible extensibility, substantially enhancing semantic accuracy and visual fidelity. User studies and expert evaluations demonstrate that the system significantly outperforms manual design in both modeling efficiency and output quality, effectively overcoming the bottlenecks of conventional monolithic models in domain adaptation and controllable generation.