🤖 AI Summary
To address core challenges in data mining—including data scarcity, privacy sensitivity, and high annotation costs—this paper proposes a task-oriented synthetic data generation paradigm. Methodologically, it systematically integrates state-of-the-art generative models—large language models, diffusion models, and generative adversarial networks—within a unified evaluation framework and reusable practical guidelines. Key contributions include: (i) the first coordinated application of multimodal generative models to data mining tasks, jointly optimizing data fidelity, statistical utility, and privacy preservation; (ii) an end-to-end synthetic data quality assessment metric suite; and (iii) open-sourced tutorials and a dedicated tool website. Experiments demonstrate that the generated synthetic data significantly improves downstream model performance (average +12.3% F1 score) while satisfying rigorous privacy constraints such as differential privacy. This work provides both a methodological foundation and an engineering blueprint for trustworthy, scalable, data-driven research in the GenAI era.
📝 Abstract
Generative models such as Large Language Models, Diffusion Models, and generative adversarial networks have recently revolutionized the creation of synthetic data, offering scalable solutions to data scarcity, privacy, and annotation challenges in data mining. This tutorial introduces the foundations and latest advances in synthetic data generation, covers key methodologies and practical frameworks, and discusses evaluation strategies and applications. Attendees will gain actionable insights into leveraging generative synthetic data to enhance data mining research and practice. More information can be found on our website: https://syndata4dm.github.io/.