Generation of Synthetic Clinical Text: A Systematic Review

📅 2025-07-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This survey systematically examines generative synthetic clinical free-text research since 2018, addressing three core questions: generation objectives (e.g., data augmentation, privacy preservation, corpus construction), technical approaches, and evaluation methodologies. Through multi-source retrieval from PubMed, ScienceDirect, arXiv, and other repositories, we identified and analyzed 94 peer-reviewed studies, conducting the first multi-dimensional quantitative analysis in this domain. Results show that Transformer-based architectures—particularly GPT-family models—have become the dominant paradigm. Synthetic clinical text demonstrates moderate fidelity as a proxy for real data in downstream NLP tasks, effectively mitigating data sparsity and improving model performance. However, privacy leakage risks remain substantial, underscoring the need for a practical–privacy co-evaluation framework integrating automated metrics with expert human review. This work fills a critical gap by providing the first systematic, quantitative review of clinical text generation, offering empirical insights to guide future research and clinical deployment.

Technology Category

Natural Language Processing: GenerationMachine Learning: PrivacyComputer Vision: Generative Adversarial Networks (GANs) for Vision

Application Category

Social Networks and Social Media: Generative AI / large language models and their impact on social systemsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Web data generation and simulation
📝 Abstract
Generating clinical synthetic text represents an effective solution for common clinical NLP issues like sparsity and privacy. This paper aims to conduct a systematic review on generating synthetic medical free-text by formulating quantitative analysis to three research questions concerning (i) the purpose of generation, (ii) the techniques, and (iii) the evaluation methods. We searched PubMed, ScienceDirect, Web of Science, Scopus, IEEE, Google Scholar, and arXiv databases for publications associated with generating synthetic medical unstructured free-text. We have identified 94 relevant articles out of 1,398 collected ones. A great deal of attention has been given to the generation of synthetic medical text from 2018 onwards, where the main purpose of such a generation is towards text augmentation, assistive writing, corpus building, privacy-preserving, annotation, and usefulness. Transformer architectures were the main predominant technique used to generate the text, especially the GPTs. On the other hand, there were four main aspects of evaluation, including similarity, privacy, structure, and utility, where utility was the most frequent method used to assess the generated synthetic medical text. Although the generated synthetic medical text demonstrated a moderate possibility to act as real medical documents in different downstream NLP tasks, it has proven to be a great asset as augmented, complementary to the real documents, towards improving the accuracy and overcoming sparsity/undersampling issues. Yet, privacy is still a major issue behind generating synthetic medical text, where more human assessments are needed to check for the existence of any sensitive information. Despite that, advances in generating synthetic medical text will considerably accelerate the adoption of workflows and pipeline development, discarding the time-consuming legalities of data transfer.
Problem

Research questions and friction points this paper is trying to address.

Systematically reviews synthetic clinical text generation techniques
Evaluates methods for privacy, utility, and similarity in synthetic text
Addresses NLP challenges like data sparsity and privacy issues
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformer architectures for synthetic text generation
Evaluation focuses on utility and privacy
Synthetic text aids NLP tasks and privacy
🔎 Similar Papers
No similar papers found.
B
Basel Alshaikhdeeb
Luxembourg Centre for Systems Biomedicine, University of Luxembourg, 6, avenue du Swing, Esch-sur-Alzette, L-4367, Luxembourg
Ahmed Abdelmonem Hemedan
Ahmed Abdelmonem Hemedan
Luxembourg University, Luxembourg Centre for Systems Biomedicine,
Computational Biology and Bioinformatics
Soumyabrata Ghosh
Soumyabrata Ghosh
University of Luxembourg
Clinical informaticsData governanceData science
I
Irina Balaur
Luxembourg Centre for Systems Biomedicine, University of Luxembourg, 6, avenue du Swing, Esch-sur-Alzette, L-4367, Luxembourg
V
Venkata Satagopam
Luxembourg Centre for Systems Biomedicine, University of Luxembourg, 6, avenue du Swing, Esch-sur-Alzette, L-4367, Luxembourg