Synthetic Data Generation Using Large Language Models: Advances in Text and Code

📅 2025-03-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the scarcity, sensitivity, and quality unreliability of labeled data in natural language and code domains by proposing a large language model (LLM)-based synthetic data generation framework. Methodologically, it introduces the first systematic integration of retrieval-augmented generation (RAG), iterative self-refinement, and execution-feedback-driven reinforcement learning from human feedback (RLHF), augmented with functional correctness verification, controllable diversity mechanisms, bias mitigation strategies, and output-weighted filtering. The key contribution is a novel synthetic data paradigm that jointly ensures accuracy, stylistic authenticity, and fairness. Extensive evaluations across classification, question answering, instruction tuning, code translation, and bug repair tasks demonstrate that models trained on synthetic data achieve performance comparable to—or even surpassing—that of models trained on real human-annotated data. Moreover, the framework substantially reduces annotation costs while preserving diversity and enabling fine-grained control over synthetic output properties.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Large Multimodal Models (LMMs)Humans and AI: Human-in-the-loop Machine Learning

Application Category

Search and Retrieval-Augmented AI: Large language models for searchEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Large language models (LLMs) have unlocked new possibilities for generating synthetic training data in both natural language and code. By producing artificial but task-relevant examples, these models can significantly augment or even replace real-world datasets, especially when labeled data is scarce or sensitive. This paper surveys recent advances in using LLMs to create synthetic text and code, emphasizing prompt-based generation, retrieval-augmented pipelines, and iterative self-refinement. We show how these methods enrich low-resource tasks such as classification and question answering, as well as code-centric applications such as instruction tuning, code translation, and bug repair, by enabling automated verification of functional correctness. Alongside potential benefits like cost-effectiveness, broad coverage, and controllable diversity, we address challenges such as factual inaccuracies in generated text, lack of stylistic realism, and the risk of bias amplification. Proposed mitigations include filtering and weighting outputs and reinforcement learning with execution feedback for code. We conclude with open research directions like automated prompt engineering, cross-modal data synthesis, and robust evaluation frameworks, highlighting the importance of LLM-generated synthetic data in advancing AI while emphasizing ethical and quality safeguards.
Problem

Research questions and friction points this paper is trying to address.

Generating synthetic training data using large language models.
Addressing scarcity and sensitivity of labeled real-world datasets.
Enhancing low-resource tasks and code-centric applications with synthetic data.
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLMs generate synthetic text and code data.
Prompt-based generation enhances low-resource tasks.
Automated verification ensures functional correctness.
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mihai Nadas
Babes ,-Bolyai University
L
Laura Dioşan
Babes ,-Bolyai University
A
Andreea Tomescu
KlusAI Labs