Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era

📅 2025-08-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address core challenges in data mining—including data scarcity, privacy sensitivity, and high annotation costs—this paper proposes a task-oriented synthetic data generation paradigm. Methodologically, it systematically integrates state-of-the-art generative models—large language models, diffusion models, and generative adversarial networks—within a unified evaluation framework and reusable practical guidelines. Key contributions include: (i) the first coordinated application of multimodal generative models to data mining tasks, jointly optimizing data fidelity, statistical utility, and privacy preservation; (ii) an end-to-end synthetic data quality assessment metric suite; and (iii) open-sourced tutorials and a dedicated tool website. Experiments demonstrate that the generated synthetic data significantly improves downstream model performance (average +12.3% F1 score) while satisfying rigorous privacy constraints such as differential privacy. This work provides both a methodological foundation and an engineering blueprint for trustworthy, scalable, data-driven research in the GenAI era.

Technology Category

Data Mining & Knowledge Management: Representing, Reasoning, and Using Provenance, TrustNatural Language Processing: GenerationMachine Learning: Privacy

Application Category

Web Mining and Content Analysis: Web data generation and simulationEconomics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAISecurity and Privacy: Data transparency and provenance
📝 Abstract
Generative models such as Large Language Models, Diffusion Models, and generative adversarial networks have recently revolutionized the creation of synthetic data, offering scalable solutions to data scarcity, privacy, and annotation challenges in data mining. This tutorial introduces the foundations and latest advances in synthetic data generation, covers key methodologies and practical frameworks, and discusses evaluation strategies and applications. Attendees will gain actionable insights into leveraging generative synthetic data to enhance data mining research and practice. More information can be found on our website: https://syndata4dm.github.io/.
Problem

Research questions and friction points this paper is trying to address.

Addressing data scarcity in data mining with synthetic data
Overcoming privacy challenges through generative synthetic data
Providing scalable annotation solutions using generative models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative models create synthetic data
Addresses data scarcity and privacy
Covers methodologies and evaluation strategies
🔎 Similar Papers
D
Dawei Li
Arizona State University, Tempe, Arizona, USA
Y
Yue Huang
University of Notre Dame, South Bend, Indiana, USA
M
Ming Li
University of Maryland, College Park, Maryland, USA
T
Tianyi Zhou
University of Maryland, College Park, Maryland, USA
Xiangliang Zhang
Xiangliang Zhang
Leonard C. Bettex Collegiate Professor, Computer Science and Engineering, University of Notre Dame
Machine LearningAI for Science
H
Huan Liu
Arizona State University, Tempe, Arizona, USA