intent-driven data synthesis

Design and build systems that synthesize labeled textual examples (e.g., utterances or dialogue turns) conditioned on explicit intent or rule specifications to produce annotation-free training data. This includes automated creation and completion of rule sets, conditioning outputs on attributes like topic or style, and scaling synthetic datasets across intents for model training and evaluation.

intent-drivendatasynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work proposes an unsupervised synthetic dialogue generation framework tailored for industrial settings where human-annotated data are scarce, relying solely on intent definitions. To enhance diversity, the approach explicitly incorporates topic and stylistic attributes and introduces two novel post-processing stylization models—Univ and Exam—combined with a large language model–based discriminative filtering mechanism to improve data quality. The study reveals that stylistic diversity has a significantly greater impact on the utility of synthetic data than topic diversity, and that integrating stylistic attributes during generation outperforms post-hoc style transfer. Experimental results demonstrate that the proposed method achieves 93.3% of the performance of models trained on human-annotated data across both industrial and public benchmarks, substantially enhancing the practicality of unlabeled synthetic dialogues.

annotation-freedialogue generationintent classification

This study addresses the problem of efficient and accurate user intent classification in large language models to facilitate downstream domain-specific model routing. For the first time, it systematically compares training-free strategies—such as lightweight methods based on internal representation statistics—with training-based approaches, including linear probes and MLP classifiers, evaluating their performance across varying task difficulty, mixed-intent prompts, and adversarial inputs. The results reveal that both paradigms achieve performance saturation on simple tasks; however, training-based methods excel in fine-grained classification (e.g., distinguishing Java from Python), whereas training-free methods demonstrate superior robustness to mixed and adversarial prompts. These findings highlight fundamental differences between the two approaches in terms of accuracy, robustness, and failure modes.

intent classificationLarge Language Modelsrobustness

This work addresses the lack of systematic optimization in existing synthetic pretraining data generation, particularly concerning prompt design, generator models, and source data selection. Through large-scale controlled experiments, we systematically investigate how to efficiently rewrite web text into high-quality synthetic data, with a focus on the impact of structured output formats (e.g., tables, math problems, FAQs), generator scale, and source data choices. Our findings reveal that structured formats substantially outperform current approaches, and that generator performance saturates beyond 1B parameters. Leveraging these insights, we propose a cost-effective synthesis strategy. Based on trillion-token-scale experiments, we release an open-source generation framework and the FinePhrase dataset—comprising 486 billion tokens—that surpasses all existing synthetic baselines in performance while reducing generation costs by up to 30×.

generator modelpretraining dataprompt design

How to Synthesize Text Data without Model Collapse?

Dec 19, 2024
XZ
Xuekai Zhu
🏛️ Shanghai Jiao Tong University | BIGAI | Peking University | Tsinghua University | Shanghai Artificial Intelligence Laboratory

This work addresses “model collapse”—the progressive degradation in performance observed when language models are iteratively trained on synthetic text data—by identifying a negative correlation between synthetic data proportion and model performance, alongside n-gram over-concentration. We propose a token-level editing method grounded in human-written text and provide the first theoretical proof that this strategy strictly bounds test error, thereby provably preventing collapse. Furthermore, we introduce a distribution-aware semi-synthetic data paradigm that overcomes the inherent degeneration bottleneck of purely generative data. Through multi-stage pretraining and fine-tuning experiments across diverse downstream tasks, our approach significantly mitigates model collapse, yielding up to a 3.2% absolute accuracy improvement. Empirical results demonstrate the superiority and robustness of semi-synthetic data over fully synthetic alternatives.

Impact of synthetic data on language model trainingPreventing model collapse in synthetic data synthesisToken-level editing to improve model performance

Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

Sep 30, 2024
KL
Ke-Han Lu
🏛️ National Taiwan University | NVIDIA

Speech-language models struggle to jointly optimize speech understanding and textual capability preservation when labeled speech instruction data is scarce. Method: We propose an end-to-end paradigm requiring zero speech instruction data. It aligns a pre-trained speech model with a large language model (LLM) via cross-modal alignment to automatically synthesize high-quality speech-text pairs, thereby injecting paralinguistic understanding; concurrently, the LLM’s parameters are frozen, and only a lightweight speech adapter is introduced to prevent textual capability degradation. Contribution/Results: This work pioneers speech-language co-modeling without any speech instruction fine-tuning data, while supporting joint optimization of complex textual instructions (e.g., chain-of-thought reasoning, format control) and speech understanding. Our approach achieves state-of-the-art performance on Dynamic-SUPERB and AIR-Bench-Chat, significantly reducing reliance on manual speech annotation.

Catastrophic ForgettingSpeech Language ModelingUnsupervised Learning

Latest Papers

What's happening recently
View more

Towards Active Synthetic Data Generation for Finetuning Language Models

Nov 30, 2025
SK
Samuel Kessler
🏛️ Microsoft | Microsoft Research NYC

This work addresses the limited adaptability of static synthetic data generation in language model fine-tuning. We propose a dynamic closed-loop synthetic data generation paradigm: during training, samples generated by a teacher model are actively selected based on the student model’s current state—such as prediction uncertainty and hidden-layer activations—enabling iterative optimization via “generate–evaluate–select–fine-tune”. Our key contribution is a lightweight, interpretable active selection strategy that significantly outperforms complex sampling methods. Evaluated on four mathematical and logical reasoning benchmarks, our approach consistently improves the performance of four small language models under fixed computational budgets, yielding average accuracy gains of 3.2–5.7 percentage points. These results demonstrate the method’s effectiveness, generalizability across diverse models and tasks, and computational efficiency.

Compares iterative versus static synthetic data generation methodsEvaluates active learning criteria for selecting synthetic training samplesOptimizes synthetic data generation for language model finetuning

To address the challenge of simultaneously achieving high quality and diversity in synthetic text data, this paper proposes GenText—a novel data synthesis framework that models semantic attributes as “textual genes” and leverages large language models (LLMs) to simulate genetic operations (crossover and mutation). Methodologically, GenText innovatively integrates genetic algorithms, attribute-driven text generation, and active learning—where the latter dynamically selects high-informativeness parent samples to guide efficient exploration of attribute combinations. Compared to conventional prompt engineering or resampling approaches, GenText significantly improves downstream model performance across multiple NLP tasks—including text classification and named entity recognition—with F1-score gains of 3.2–5.8 percentage points under class-imbalanced settings. The framework’s source code and datasets are publicly released.

Enhancing synthetic data quality and diversity using genetic algorithmsOptimizing parent selection to expand offspring search spaceSimulating genetic operations with LLMs for novel attribute combinations

To address the scarcity of paired input-output data in low-resource natural language generation (NLG), this paper proposes PbT, a two-stage teacher-student framework. The teacher model compresses unpaired inputs and outputs separately into compact, shared intermediate representations; the student model then learns to reconstruct the original inputs from these representations, thereby synthesizing high-fidelity pseudo-paired data. Crucially, PbT bridges unpaired data via intermediate representations—eliminating reliance on costly human annotation or direct large-model generation, which suffers from high computational expense and poor generalization. Evaluated on five benchmarks, an 8B student model trained solely on PbT-synthesized data achieves a ROUGE-L score significantly surpassing that obtained using data generated by a 70B model, and approaches human-annotated performance—narrowing the gap to just 1.2 points and closing 82% of the oracle gap—while reducing annotation cost to one-third that of direct synthesis.

Addressing data scarcity in low-resource text generationCreating high-fidelity synthetic training data efficientlyGenerating input-output pairs without parallel data

This work addresses the limitation of existing large language models in synthetic data generation, which typically treat tasks as isolated events and thus fail to accumulate or transfer synthesis experience across tasks. To overcome this, the authors propose StreamSynth, a novel paradigm that formulates synthetic data generation as an experience-driven continual learning process. By incorporating streaming task inputs and a feedback mechanism, StreamSynth enables the model to continuously learn from and reuse effective synthesis strategies across a sequence of tasks. The proposed SynLearner framework integrates diverse exploration, feedback-based learning, and a balanced optimization of quality and diversity. Experimental results demonstrate that the approach effectively leverages early-task experience to enhance performance on subsequent tasks, exhibiting robust cross-task transfer and cumulative learning capabilities across multiple benchmarks.

cross-task transferexperience accumulationlarge language models

This work addresses the challenge that current large language models struggle to align with user intent during pretraining due to insufficient supervised instruction data. To overcome this limitation, we introduce FineInstructions, a large-scale synthetic dataset comprising billions of high-quality instruction–response pairs, automatically generated by matching internet-scale unstructured corpora with approximately 18 million instruction templates derived from real user queries. Leveraging this dataset, we present the first approach to pretrain a language model from scratch using purely instruction-tuning objectives, thereby departing from conventional self-supervised paradigms. Experimental results demonstrate that, at equal token budgets, our method significantly outperforms standard pretraining and alternative synthetic data strategies, achieving superior response quality on standard benchmarks for open-ended generation tasks.

instruction tuninglarge language modelspre-training

Hot Scholars

YC

Yoonseo Choi

KAIST
Human-AI InteractionLLM SimulationGenerative AgentDesign Process
JC

Junjie Chen

Professor, Tianjin University
Software TestingCompiler TestingAI4SESE4AI
IK

Ishita Khan

Search Science, eBay Inc
Applied Machine LearningData ScienceSearchInformation Retrieval
SN

Sophie Nagler

University of Oxford
Mathematical PhilosophyLogicPhilosophy of SciencePhilosophy of Mathematics