Generative adversarial networks vs large language models: a comparative study on synthetic tabular data generation

📅 2025-02-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the feasibility and advantages of leveraging large language models (LLMs) for zero-shot high-fidelity synthetic tabular data generation, challenging the dominance of generative adversarial networks (GANs). We propose a zero-shot tabular generation framework based on GPT-4o and natural-language prompting—requiring no fine-tuning, pretraining, or exposure to real data. To our knowledge, this is the first systematic comparison of LLMs and CTGAN across three critical dimensions: statistical fidelity, multivariate dependency preservation, and privacy protection. Evaluations on three benchmark datasets—including Iris—demonstrate that GPT-4o significantly outperforms CTGAN in estimating means and confidence intervals, preserving both direction and strength of bivariate correlations, and inherently eliminating training-data leakage risks. It exhibits only minor limitations in modeling marginal distribution shapes. These results establish LLMs as a lightweight, privacy-preserving, and plug-and-play paradigm for synthetic tabular data generation.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: GenerationComputer Vision: Generative Adversarial Networks (GANs) for Vision

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSocial Networks and Social Media: Generative AI / large language models and their impact on social systemsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
We propose a new framework for zero-shot generation of synthetic tabular data. Using the large language model (LLM) GPT-4o and plain-language prompting, we demonstrate the ability to generate high-fidelity tabular data without task-specific fine-tuning or access to real-world data (RWD) for pre-training. To benchmark GPT-4o, we compared the fidelity and privacy of LLM-generated synthetic data against data generated with the conditional tabular generative adversarial network (CTGAN), across three open-access datasets: Iris, Fish Measurements, and Real Estate Valuation. Despite the zero-shot approach, GPT-4o outperformed CTGAN in preserving means, 95% confidence intervals, bivariate correlations, and data privacy of RWD, even at amplified sample sizes. Notably, correlations between parameters were consistently preserved with appropriate direction and strength. However, refinement is necessary to better retain distributional characteristics. These findings highlight the potential of LLMs in tabular data synthesis, offering an accessible alternative to generative adversarial networks and variational autoencoders.
Problem

Research questions and friction points this paper is trying to address.

Compare LLM and GAN for tabular data generation.
Propose zero-shot synthetic data generation framework.
Assess fidelity and privacy of synthetic data.
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM for zero-shot data generation
GPT-4o outperforms CTGAN benchmarks
Preserves data privacy and correlations
A
Austin A. Barr
Cumming School of Medicine, University of Calgary, Calgary, AB, Canada
R
Robert Rozman
Independent Researcher, Toronto, ON, Canada
Eddie Guo
Eddie Guo
University of Toronto
Large Language ModelsMachine LearningMedical EducationNeurosurgery