Zero-shot generation of synthetic neurosurgical data with large language models

📅 2025-02-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Neurosurgery faces critical bottlenecks in real-world data (RWD) acquisition—including scarcity of high-quality samples, stringent privacy regulations, and high preprocessing costs. To address these challenges, this work proposes, for the first time, a zero-shot large language model (LLM)-based paradigm for synthesizing neurosurgical data using GPT-4o—requiring no RWD for training or fine-tuning and relying solely on prompt engineering to generate high-fidelity synthetic data. We systematically benchmark against CTGAN across statistical fidelity (marginal means, distributions, pairwise correlations), machine learning utility (achieving F1 = 0.706 in prognostic prediction), and privacy preservation (near-zero record duplication rate). Results demonstrate that GPT-4o outperforms CTGAN in univariate/bivariate fidelity and classification performance, while enabling robust modeling under small-sample regimes. This study pioneers an LLM-driven, zero-contact, high-fidelity, and privacy-preserving synthetic data generation framework for neurosurgical RWD.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: PrivacyComputer Vision: Generative Adversarial Networks (GANs) for Vision

Application Category

Economics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAISocial Networks and Social Media: Generative AI / large language models and their impact on social systemsSecurity and Privacy: Data transparency and provenance
📝 Abstract
Clinical data is fundamental to advance neurosurgical research, but access is often constrained by data availability, small sample sizes, privacy regulations, and resource-intensive preprocessing and de-identification procedures. Synthetic data offers a potential solution to challenges associated with accessing and using real-world data (RWD). This study aims to evaluate the capability of zero-shot generation of synthetic neurosurgical data with a large language model (LLM), GPT-4o, by benchmarking with the conditional tabular generative adversarial network (CTGAN). Synthetic datasets were compared to real-world neurosurgical data to assess fidelity (means, proportions, distributions, and bivariate correlations), utility (ML classifier performance on RWD), and privacy (duplication of records from RWD). The GPT-4o-generated datasets matched or exceeded CTGAN performance, despite no fine-tuning or access to RWD for pre-training. Datasets demonstrated high univariate and bivariate fidelity to RWD without directly exposing any real patient records, even at amplified sample size. Training an ML classifier on GPT-4o-generated data and testing on RWD for a binary prediction task showed an F1 score (0.706) with comparable performance to training on the CTGAN data (0.705) for predicting postoperative functional status deterioration. GPT-4o demonstrated a promising ability to generate high-fidelity synthetic neurosurgical data. These findings also indicate that data synthesized with GPT-4o can effectively augment clinical data with small sample sizes, and train ML models for prediction of neurosurgical outcomes. Further investigation is necessary to improve the preservation of distributional characteristics and boost classifier performance.
Problem

Research questions and friction points this paper is trying to address.

Generate synthetic neurosurgical data using GPT-4.
Evaluate fidelity, utility, and privacy of synthetic data.
Augment small clinical datasets with synthetic data.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-shot synthetic data generation
GPT-4o outperforms CTGAN
High-fidelity neurosurgical data synthesis
🔎 Similar Papers
2024-07-12Pattern Recognition LettersCitations: 2