Score
Designs and implements pipelines and methods that use large language models to generate synthetic textual datasets and labeled example pairs, including paraphrases, LLM-style outputs, and other synthetic examples. Evaluates and calibrates how such LLM-produced data affects downstream training, detectors, classifiers, and distribution shifts between pre- and post-LLM text.
In low-resource settings, large language models (LLMs) suffer from insufficient training data, while existing text data augmentation methods struggle to simultaneously ensure enhanced text quality and factual fidelity. Method: This paper systematically surveys LLM-driven text data augmentation techniques and proposes a unified taxonomy—comprising simple, prompt-based, retrieval-augmented, and hybrid augmentation—alongside a novel “generate–retrieve–verify” collaborative augmentation paradigm that emphasizes post-hoc verification for improving data trustworthiness. It integrates prompt engineering, external knowledge retrieval, generation-quality filtering, and multi-dimensional evaluation to enable task-controllable, high-fidelity text generation. Contribution/Results: The work delivers a comprehensive technical landscape covering methodologies, open challenges, and emerging opportunities, yielding a reusable, empirically verifiable augmentation framework specifically designed to advance low-resource LLM training.
High-quality annotated data for non-English languages—such as Italian—is scarce and expensive to collect, hindering inclusive language detection in domains like job advertisements. Method: This work introduces an end-to-end synthetic data generation and evaluation framework tailored to Italian recruitment texts. It systematically investigates the impact of prompting strategies, text length, and target position on LLM-generated data quality—the first such study for this task. We propose multi-dimensional controllable prompt engineering, fine-grained inclusive language modeling, and cross-distribution evaluation. Results: Models fine-tuned exclusively on synthetic data significantly outperform baselines on both real-world and synthetic test sets, demonstrating strong generalization and robustness. Our framework establishes a reproducible, scalable, and data-efficient paradigm for LLM adaptation in low-resource languages, advancing inclusive NLP for under-resourced linguistic settings.
This work addresses the scarcity, sensitivity, and quality unreliability of labeled data in natural language and code domains by proposing a large language model (LLM)-based synthetic data generation framework. Methodologically, it introduces the first systematic integration of retrieval-augmented generation (RAG), iterative self-refinement, and execution-feedback-driven reinforcement learning from human feedback (RLHF), augmented with functional correctness verification, controllable diversity mechanisms, bias mitigation strategies, and output-weighted filtering. The key contribution is a novel synthetic data paradigm that jointly ensures accuracy, stylistic authenticity, and fairness. Extensive evaluations across classification, question answering, instruction tuning, code translation, and bug repair tasks demonstrate that models trained on synthetic data achieve performance comparable to—or even surpassing—that of models trained on real human-annotated data. Moreover, the framework substantially reduces annotation costs while preserving diversity and enabling fine-grained control over synthetic output properties.
Synthetic natural language descriptions generated by large language models (LLMs) are increasingly used to train spreadsheet formula generation models, yet their annotation quality—and its impact on downstream fine-tuning performance—remains poorly understood. Method: We propose a proxy-objective-based synthetic data validation framework that integrates multi-model comparison (two open-source and two closed-source LLMs) with rigorous formula–natural language alignment evaluation. Contribution/Results: Through systematic empirical analysis, we demonstrate for the first time that synthetic annotation quality critically influences fine-tuning efficacy. While high-quality sample filtering reduces dataset size, it unexpectedly enhances model generalization and reasoning capabilities on complex formulas. Our approach consistently improves fine-tuning performance across four state-of-the-art models, confirming that high-fidelity synthetic data is a key lever for improving robustness in formula generation systems.
Existing methods struggle to simultaneously achieve high accuracy in distinguishing human-written text from AI-generated text and precisely identifying the specific large language model (LLM) responsible for generation (e.g., GPT-4o-mini, LLaMA-3, BERT). Method: We propose a unified discriminative framework that jointly performs AI-text detection and LLM attribution via instruction-tuned fine-tuning of three heterogeneous models—GPT-4o-mini, LLaMA-3-8B, and BERT—within a single end-to-end pipeline. Contribution/Results: Our approach introduces cross-architecture collaborative training and tightly coupled multi-task learning, enhancing robustness under complex, real-world conditions. Experiments demonstrate 95.47% accuracy in AI-text detection and 46.98% accuracy in fine-grained LLM attribution—substantially outperforming prior baselines. This framework provides a scalable, technically grounded solution for combating disinformation and ensuring AI content traceability.
High-quality pretraining of large language models (LLMs) is constrained by the scarcity of natural, high-quality textual data, raising the critical question of whether synthetic data can effectively substitute for or augment natural data. Method: We conduct a large-scale ablation study under a unified experimental protocol—training over 1,000 models using >100,000 GPU-hours—systematically evaluating diverse synthetic data types (e.g., paraphrased text, generated textbooks) and their mixed-data training strategies with natural corpora. Contribution/Results: We find that mixing 30% paraphrased data accelerates convergence by 5–10×; synthetic data yields diminishing returns dependent on model scale; and susceptibility to “model collapse” varies significantly across synthetic data types. Crucially, we propose the first practical, model-size- and data-budget-aware heuristic for dynamically allocating synthetic-to-natural data ratios. We empirically refute pure synthetic pretraining but demonstrate that judicious hybrid training achieves both computational efficiency and training stability—providing an evidence-based, scalable paradigm for LLM pretraining.
This work addresses the challenge of limited labeled data in low-resource multilingual settings, which severely hampers the performance of small classification models. The authors propose a novel paradigm that leverages large multilingual language models as "teachers" to generate high-quality synthetic data through instruction tuning and in-context learning, enabling cross-lingual knowledge distillation into lightweight student models. Experimental results demonstrate that student models trained on only a small amount of such synthesized data consistently outperform the original large language model across 11 languages and four text classification tasks, with particularly pronounced gains in low-resource languages. These findings validate the efficacy and efficiency of employing large language models as data generators rather than direct classifiers in resource-constrained multilingual scenarios.
This work addresses the limited systematic understanding of how data characteristics influence large language model performance across training, fine-tuning, alignment, and in-context learning stages—a gap often filled by computationally expensive empirical trial-and-error. To overcome this, the paper introduces a novel “data probing” paradigm that leverages stochastic processes to generate synthetic data with controllable statistical properties. By integrating tools from information theory, such as typical sets, the authors construct an interpretable and tunable experimental framework. This approach systematically elucidates the mechanisms through which data properties affect model performance, generalization, and robustness, offering both a theoretical foundation and an efficient experimental pathway to replace conventional heuristic data selection practices.
To address the challenges of tracing provenance and ensuring credibility of large language model (LLM)–generated content, this paper proposes the first holistic four-dimensional provenance framework integrating both model- and data-centric perspectives: model origin identification, architectural and mechanistic analysis, training data attribution, and external information verification. We introduce a novel “prior–posterior” dual-paradigm classification system and unify techniques including model fingerprinting, response-level verification, and traceability-aware embedding to support both proactive and reactive reasoning. The framework systematically consolidates fragmented provenance research efforts, significantly enhancing the explainability, verifiability, and transparency of AI-generated content. It establishes a theoretical foundation and scalable technical methodology for detecting AI-generated content (AIGC), identifying model identities, and ensuring information reliability.
This study addresses the challenge of generating executable scientific code for novel algorithms using large language models (LLMs) in zero-shot, training-free settings. To overcome the limitations of existing tools like Code-Scribe—which lack support for zero-shot algorithm implementation—we propose an LLM-assisted progressive code synthesis framework. It integrates program semantic understanding, algorithmic structure parsing, and iterative code verification to enable end-to-end generation of high-fidelity scientific computing code from natural language specifications. Unlike conventional data-driven approaches, our method eliminates reliance on historical code examples and successfully automates the implementation of original numerical algorithms—including custom integrators and optimizers—without any task-specific training data. Experiments demonstrate 89.3% functional correctness and a 72% reduction in average debugging time, significantly accelerating scientific software extensibility. The results validate the feasibility and engineering utility of LLMs in creative, specification-driven programming tasks.
This work addresses the challenges of applying large language models (LLMs) in modeling and simulation (M&S), where suboptimal prompt design, improper hyperparameter configuration, or inadequate data handling often lead to performance degradation, information loss, and non-deterministic behavior. For the first time, this study systematically identifies latent pitfalls specific to LLM deployment in M&S and proposes a principled framework centered on rigorous design and empirical evaluation. The framework encompasses key techniques including prompt engineering, retrieval-augmented generation (RAG), low-rank adaptation (LoRA), temperature control, and context management. By offering a structured set of practical guidelines, this research enables practitioners to critically assess the suitability and implementation strategies of LLMs in M&S contexts, thereby substantially enhancing their effectiveness and reliability.