llm data synthesis

Designs and implements pipelines and methods that use large language models to generate synthetic textual datasets and labeled example pairs, including paraphrases, LLM-style outputs, and other synthetic examples. Evaluates and calibrates how such LLM-produced data affects downstream training, detectors, classifiers, and distribution shifts between pre- and post-LLM text.

llmdatasynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.7
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

High-quality annotated data for non-English languages—such as Italian—is scarce and expensive to collect, hindering inclusive language detection in domains like job advertisements. Method: This work introduces an end-to-end synthetic data generation and evaluation framework tailored to Italian recruitment texts. It systematically investigates the impact of prompting strategies, text length, and target position on LLM-generated data quality—the first such study for this task. We propose multi-dimensional controllable prompt engineering, fine-grained inclusive language modeling, and cross-distribution evaluation. Results: Models fine-tuned exclusively on synthetic data significantly outperform baselines on both real-world and synthetic test sets, demonstrating strong generalization and robustness. Our framework establishes a reproducible, scalable, and data-efficient paradigm for LLM adaptation in low-resource languages, advancing inclusive NLP for under-resourced linguistic settings.

Evaluating synthetic data validity factors in LLMsGenerating synthetic data for non-English language tasksImproving inclusive language detection in Italian job ads

Synthetic Data Generation Using Large Language Models: Advances in Text and Code

Mar 18, 2025
MN
Mihai Nadas
🏛️ Babes ,-Bolyai University | KlusAI Labs

This work addresses the scarcity, sensitivity, and quality unreliability of labeled data in natural language and code domains by proposing a large language model (LLM)-based synthetic data generation framework. Methodologically, it introduces the first systematic integration of retrieval-augmented generation (RAG), iterative self-refinement, and execution-feedback-driven reinforcement learning from human feedback (RLHF), augmented with functional correctness verification, controllable diversity mechanisms, bias mitigation strategies, and output-weighted filtering. The key contribution is a novel synthetic data paradigm that jointly ensures accuracy, stylistic authenticity, and fairness. Extensive evaluations across classification, question answering, instruction tuning, code translation, and bug repair tasks demonstrate that models trained on synthetic data achieve performance comparable to—or even surpassing—that of models trained on real human-annotated data. Moreover, the framework substantially reduces annotation costs while preserving diversity and enabling fine-grained control over synthetic output properties.

Addressing scarcity and sensitivity of labeled real-world datasets.Enhancing low-resource tasks and code-centric applications with synthetic data.Generating synthetic training data using large language models.

Synthetic natural language descriptions generated by large language models (LLMs) are increasingly used to train spreadsheet formula generation models, yet their annotation quality—and its impact on downstream fine-tuning performance—remains poorly understood. Method: We propose a proxy-objective-based synthetic data validation framework that integrates multi-model comparison (two open-source and two closed-source LLMs) with rigorous formula–natural language alignment evaluation. Contribution/Results: Through systematic empirical analysis, we demonstrate for the first time that synthetic annotation quality critically influences fine-tuning efficacy. While high-quality sample filtering reduces dataset size, it unexpectedly enhances model generalization and reasoning capabilities on complex formulas. Our approach consistently improves fine-tuning performance across four state-of-the-art models, confirming that high-fidelity synthetic data is a key lever for improving robustness in formula generation systems.

Assessing impact of validation on LLM fine-tuning performanceExploring trade-off between example difficulty and model capabilityValidating synthetic NL annotations for spreadsheet formula generation

AI Generated Text Detection Using Instruction Fine-tuned Large Language and Transformer-Based Models

Jul 07, 2025
CG
Chinnappa Guggilla
🏛️ Deloitte & Touche Assurance and Enterprise Risk Services India Private Limited | Deloitte & Touche LLP

Existing methods struggle to simultaneously achieve high accuracy in distinguishing human-written text from AI-generated text and precisely identifying the specific large language model (LLM) responsible for generation (e.g., GPT-4o-mini, LLaMA-3, BERT). Method: We propose a unified discriminative framework that jointly performs AI-text detection and LLM attribution via instruction-tuned fine-tuning of three heterogeneous models—GPT-4o-mini, LLaMA-3-8B, and BERT—within a single end-to-end pipeline. Contribution/Results: Our approach introduces cross-architecture collaborative training and tightly coupled multi-task learning, enhancing robustness under complex, real-world conditions. Experiments demonstrate 95.47% accuracy in AI-text detection and 46.98% accuracy in fine-grained LLM attribution—substantially outperforming prior baselines. This framework provides a scalable, technically grounded solution for combating disinformation and ensuring AI content traceability.

Detect AI-generated text to prevent misuse in phishing and fake news.Identify specific LLM models responsible for generating the text.Improve detection accuracy using fine-tuned GPT, LLaMA, and BERT models.

Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls

Oct 01, 2025
FK
Feiyang Kang
🏛️ FAIR at Meta | Virginia Tech | Cerebras Systems | Independent consultant

High-quality pretraining of large language models (LLMs) is constrained by the scarcity of natural, high-quality textual data, raising the critical question of whether synthetic data can effectively substitute for or augment natural data. Method: We conduct a large-scale ablation study under a unified experimental protocol—training over 1,000 models using >100,000 GPU-hours—systematically evaluating diverse synthetic data types (e.g., paraphrased text, generated textbooks) and their mixed-data training strategies with natural corpora. Contribution/Results: We find that mixing 30% paraphrased data accelerates convergence by 5–10×; synthetic data yields diminishing returns dependent on model scale; and susceptibility to “model collapse” varies significantly across synthetic data types. Crucially, we propose the first practical, model-size- and data-budget-aware heuristic for dynamically allocating synthetic-to-natural data ratios. We empirically refute pure synthetic pretraining but demonstrate that judicious hybrid training achieves both computational efficiency and training stability—providing an evidence-based, scalable paradigm for LLM pretraining.

Determining optimal synthetic-natural data mixtures to accelerate training convergenceEvaluating synthetic data's effectiveness versus natural web data for LLM pre-trainingInvestigating model collapse risks across different synthetic data generation methods

Latest Papers

What's happening recently
View more

This work addresses the challenge of limited labeled data in low-resource multilingual settings, which severely hampers the performance of small classification models. The authors propose a novel paradigm that leverages large multilingual language models as "teachers" to generate high-quality synthetic data through instruction tuning and in-context learning, enabling cross-lingual knowledge distillation into lightweight student models. Experimental results demonstrate that student models trained on only a small amount of such synthesized data consistently outperform the original large language model across 11 languages and four text classification tasks, with particularly pronounced gains in low-resource languages. These findings validate the efficacy and efficiency of employing large language models as data generators rather than direct classifiers in resource-constrained multilingual scenarios.

data scarcitylarge language modelslow-resource

This work addresses the limited systematic understanding of how data characteristics influence large language model performance across training, fine-tuning, alignment, and in-context learning stages—a gap often filled by computationally expensive empirical trial-and-error. To overcome this, the paper introduces a novel “data probing” paradigm that leverages stochastic processes to generate synthetic data with controllable statistical properties. By integrating tools from information theory, such as typical sets, the authors construct an interpretable and tunable experimental framework. This approach systematically elucidates the mechanisms through which data properties affect model performance, generalization, and robustness, offering both a theoretical foundation and an efficient experimental pathway to replace conventional heuristic data selection practices.

data characteristicsdata probinglarge language models

Large Language Model Sourcing: A Survey

Oct 11, 2025
LP
Liang Pang
🏛️ State Key Laboratory of AI Safety | Institute of Computing Technology, Chinese Academy of Sciences | University of Chinese Academy of Sciences | Gaoling School of Artificial Intelligence | Renmin University of China

To address the challenges of tracing provenance and ensuring credibility of large language model (LLM)–generated content, this paper proposes the first holistic four-dimensional provenance framework integrating both model- and data-centric perspectives: model origin identification, architectural and mechanistic analysis, training data attribution, and external information verification. We introduce a novel “prior–posterior” dual-paradigm classification system and unify techniques including model fingerprinting, response-level verification, and traceability-aware embedding to support both proactive and reactive reasoning. The framework systematically consolidates fragmented provenance research efforts, significantly enhancing the explainability, verifiability, and transparency of AI-generated content. It establishes a theoretical foundation and scalable technical methodology for detecting AI-generated content (AIGC), identifying model identities, and ensuring information reliability.

Addressing hallucinations and bias through multi-perspective sourcingDeveloping traceability methods for model structure and training dataTracking provenance of LLM-generated content to enhance transparency

Adding New Capability in Existing Scientific Application with LLM Assistance

Oct 29, 2025
AD
Anshu Dubey
🏛️ Argonne National Laboratory

This study addresses the challenge of generating executable scientific code for novel algorithms using large language models (LLMs) in zero-shot, training-free settings. To overcome the limitations of existing tools like Code-Scribe—which lack support for zero-shot algorithm implementation—we propose an LLM-assisted progressive code synthesis framework. It integrates program semantic understanding, algorithmic structure parsing, and iterative code verification to enable end-to-end generation of high-fidelity scientific computing code from natural language specifications. Unlike conventional data-driven approaches, our method eliminates reliance on historical code examples and successfully automates the implementation of original numerical algorithms—including custom integrators and optimizers—without any task-specific training data. Experiments demonstrate 89.3% functional correctness and a 72% reduction in average debugging time, significantly accelerating scientific software extensibility. The results validate the feasibility and engineering utility of LLMs in creative, specification-driven programming tasks.

Automating coding tasks using large language modelsEnhancing code-translation tools for new code generationGenerating code for new algorithms without training examples

This work addresses the challenges of applying large language models (LLMs) in modeling and simulation (M&S), where suboptimal prompt design, improper hyperparameter configuration, or inadequate data handling often lead to performance degradation, information loss, and non-deterministic behavior. For the first time, this study systematically identifies latent pitfalls specific to LLM deployment in M&S and proposes a principled framework centered on rigorous design and empirical evaluation. The framework encompasses key techniques including prompt engineering, retrieval-augmented generation (RAG), low-rank adaptation (LoRA), temperature control, and context management. By offering a structured set of practical guidelines, this research enables practitioners to critically assess the suitability and implementation strategies of LLMs in M&S contexts, thereby substantially enhancing their effectiveness and reliability.

Hyper-parameter TuningLarge Language ModelsModeling and Simulation

Hot Scholars

JL

Jian Luan

Toshiba, Microsoft, Xiaomi
LLMVLMTTSSinging Synthesis
YS

Yuanfeng Song

Unknown affiliation
NLP4DataData VisualizationText2SQLLLM
DL

Dongha Lee

Yonsei University
Data miningInformation retrievalNatural language processing
HY

Hwanjo Yu

POSTECH
data miningmachine learningrecommendation systemtime-series
SK

SeongKu Kang

Korea University
Data miningText miningRecommender SystemInformation Retrieval