Score
Design and build systems that synthesize labeled textual examples (e.g., utterances or dialogue turns) conditioned on explicit intent or rule specifications to produce annotation-free training data. This includes automated creation and completion of rule sets, conditioning outputs on attributes like topic or style, and scaling synthetic datasets across intents for model training and evaluation.
In low-resource settings, large language models (LLMs) suffer from insufficient training data, while existing text data augmentation methods struggle to simultaneously ensure enhanced text quality and factual fidelity. Method: This paper systematically surveys LLM-driven text data augmentation techniques and proposes a unified taxonomy—comprising simple, prompt-based, retrieval-augmented, and hybrid augmentation—alongside a novel “generate–retrieve–verify” collaborative augmentation paradigm that emphasizes post-hoc verification for improving data trustworthiness. It integrates prompt engineering, external knowledge retrieval, generation-quality filtering, and multi-dimensional evaluation to enable task-controllable, high-fidelity text generation. Contribution/Results: The work delivers a comprehensive technical landscape covering methodologies, open challenges, and emerging opportunities, yielding a reusable, empirically verifiable augmentation framework specifically designed to advance low-resource LLM training.
This work proposes an unsupervised synthetic dialogue generation framework tailored for industrial settings where human-annotated data are scarce, relying solely on intent definitions. To enhance diversity, the approach explicitly incorporates topic and stylistic attributes and introduces two novel post-processing stylization models—Univ and Exam—combined with a large language model–based discriminative filtering mechanism to improve data quality. The study reveals that stylistic diversity has a significantly greater impact on the utility of synthetic data than topic diversity, and that integrating stylistic attributes during generation outperforms post-hoc style transfer. Experimental results demonstrate that the proposed method achieves 93.3% of the performance of models trained on human-annotated data across both industrial and public benchmarks, substantially enhancing the practicality of unlabeled synthetic dialogues.
This study addresses the problem of efficient and accurate user intent classification in large language models to facilitate downstream domain-specific model routing. For the first time, it systematically compares training-free strategies—such as lightweight methods based on internal representation statistics—with training-based approaches, including linear probes and MLP classifiers, evaluating their performance across varying task difficulty, mixed-intent prompts, and adversarial inputs. The results reveal that both paradigms achieve performance saturation on simple tasks; however, training-based methods excel in fine-grained classification (e.g., distinguishing Java from Python), whereas training-free methods demonstrate superior robustness to mixed and adversarial prompts. These findings highlight fundamental differences between the two approaches in terms of accuracy, robustness, and failure modes.
This work addresses the lack of systematic optimization in existing synthetic pretraining data generation, particularly concerning prompt design, generator models, and source data selection. Through large-scale controlled experiments, we systematically investigate how to efficiently rewrite web text into high-quality synthetic data, with a focus on the impact of structured output formats (e.g., tables, math problems, FAQs), generator scale, and source data choices. Our findings reveal that structured formats substantially outperform current approaches, and that generator performance saturates beyond 1B parameters. Leveraging these insights, we propose a cost-effective synthesis strategy. Based on trillion-token-scale experiments, we release an open-source generation framework and the FinePhrase dataset—comprising 486 billion tokens—that surpasses all existing synthetic baselines in performance while reducing generation costs by up to 30×.
This work addresses “model collapse”—the progressive degradation in performance observed when language models are iteratively trained on synthetic text data—by identifying a negative correlation between synthetic data proportion and model performance, alongside n-gram over-concentration. We propose a token-level editing method grounded in human-written text and provide the first theoretical proof that this strategy strictly bounds test error, thereby provably preventing collapse. Furthermore, we introduce a distribution-aware semi-synthetic data paradigm that overcomes the inherent degeneration bottleneck of purely generative data. Through multi-stage pretraining and fine-tuning experiments across diverse downstream tasks, our approach significantly mitigates model collapse, yielding up to a 3.2% absolute accuracy improvement. Empirical results demonstrate the superiority and robustness of semi-synthetic data over fully synthetic alternatives.
Speech-language models struggle to jointly optimize speech understanding and textual capability preservation when labeled speech instruction data is scarce. Method: We propose an end-to-end paradigm requiring zero speech instruction data. It aligns a pre-trained speech model with a large language model (LLM) via cross-modal alignment to automatically synthesize high-quality speech-text pairs, thereby injecting paralinguistic understanding; concurrently, the LLM’s parameters are frozen, and only a lightweight speech adapter is introduced to prevent textual capability degradation. Contribution/Results: This work pioneers speech-language co-modeling without any speech instruction fine-tuning data, while supporting joint optimization of complex textual instructions (e.g., chain-of-thought reasoning, format control) and speech understanding. Our approach achieves state-of-the-art performance on Dynamic-SUPERB and AIR-Bench-Chat, significantly reducing reliance on manual speech annotation.
This work addresses the limited adaptability of static synthetic data generation in language model fine-tuning. We propose a dynamic closed-loop synthetic data generation paradigm: during training, samples generated by a teacher model are actively selected based on the student model’s current state—such as prediction uncertainty and hidden-layer activations—enabling iterative optimization via “generate–evaluate–select–fine-tune”. Our key contribution is a lightweight, interpretable active selection strategy that significantly outperforms complex sampling methods. Evaluated on four mathematical and logical reasoning benchmarks, our approach consistently improves the performance of four small language models under fixed computational budgets, yielding average accuracy gains of 3.2–5.7 percentage points. These results demonstrate the method’s effectiveness, generalizability across diverse models and tasks, and computational efficiency.
To address the challenge of simultaneously achieving high quality and diversity in synthetic text data, this paper proposes GenText—a novel data synthesis framework that models semantic attributes as “textual genes” and leverages large language models (LLMs) to simulate genetic operations (crossover and mutation). Methodologically, GenText innovatively integrates genetic algorithms, attribute-driven text generation, and active learning—where the latter dynamically selects high-informativeness parent samples to guide efficient exploration of attribute combinations. Compared to conventional prompt engineering or resampling approaches, GenText significantly improves downstream model performance across multiple NLP tasks—including text classification and named entity recognition—with F1-score gains of 3.2–5.8 percentage points under class-imbalanced settings. The framework’s source code and datasets are publicly released.
To address the scarcity of paired input-output data in low-resource natural language generation (NLG), this paper proposes PbT, a two-stage teacher-student framework. The teacher model compresses unpaired inputs and outputs separately into compact, shared intermediate representations; the student model then learns to reconstruct the original inputs from these representations, thereby synthesizing high-fidelity pseudo-paired data. Crucially, PbT bridges unpaired data via intermediate representations—eliminating reliance on costly human annotation or direct large-model generation, which suffers from high computational expense and poor generalization. Evaluated on five benchmarks, an 8B student model trained solely on PbT-synthesized data achieves a ROUGE-L score significantly surpassing that obtained using data generated by a 70B model, and approaches human-annotated performance—narrowing the gap to just 1.2 points and closing 82% of the oracle gap—while reducing annotation cost to one-third that of direct synthesis.
This work addresses the limitation of existing large language models in synthetic data generation, which typically treat tasks as isolated events and thus fail to accumulate or transfer synthesis experience across tasks. To overcome this, the authors propose StreamSynth, a novel paradigm that formulates synthetic data generation as an experience-driven continual learning process. By incorporating streaming task inputs and a feedback mechanism, StreamSynth enables the model to continuously learn from and reuse effective synthesis strategies across a sequence of tasks. The proposed SynLearner framework integrates diverse exploration, feedback-based learning, and a balanced optimization of quality and diversity. Experimental results demonstrate that the approach effectively leverages early-task experience to enhance performance on subsequent tasks, exhibiting robust cross-task transfer and cumulative learning capabilities across multiple benchmarks.
This work addresses the challenge that current large language models struggle to align with user intent during pretraining due to insufficient supervised instruction data. To overcome this limitation, we introduce FineInstructions, a large-scale synthetic dataset comprising billions of high-quality instruction–response pairs, automatically generated by matching internet-scale unstructured corpora with approximately 18 million instruction templates derived from real user queries. Leveraging this dataset, we present the first approach to pretrain a language model from scratch using purely instruction-tuning objectives, thereby departing from conventional self-supervised paradigms. Experimental results demonstrate that, at equal token budgets, our method significantly outperforms standard pretraining and alternative synthetic data strategies, achieving superior response quality on standard benchmarks for open-ended generation tasks.