instruction dataset generation

Designs, curates, and synthesizes datasets of instruction-style examples—covering instruction–input–output triples, instruction–code pairs, and instruction-linked knowledge items—including construction and maintenance of instruction knowledge bases and task-specific corpora. Defines instruction templates, parsing and encoding schemes, generates synthetic instructions and multi-stage prompt sequences, prepares data for instruction tuning/fine-tuning, and builds evaluation protocols and benchmarks to measure instruction-following behavior.

instructiondatasetgeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$130K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Dynamics of Instruction Tuning: Each Ability of Large Language Models Has Its Own Growth Pace

Oct 30, 2023
CS
Chiyu Song
🏛️ Zhejiang University | Westlake University | Westlake Institute for Advanced Study

This work investigates the mechanisms by which instruction tuning enhances general intelligence in Chinese large language models (LLMs), focusing on how data scale, model size (7B–33B), and data construction methodology (human-authored vs. synthetic) differentially affect multidimensional capabilities—including creative writing, code generation, and logical reasoning. Method: Leveraging a 40k+ multi-capability-annotated instruction dataset, we conduct cross-domain ablation studies to isolate these factors. Contribution/Results: We first reveal that underlying capabilities evolve at independent learning paces; human-authored data remains consistently effective, whereas synthetic data exhibits a performance ceiling; and instruction data demonstrates strong cross-capability generalization. Based on these findings, we propose a quantifiable, efficiency-oriented data construction guideline. Evaluated on two public benchmarks, our approach yields significant performance gains, providing empirical evidence and methodological foundations for capability-targeted LLM optimization.

Explores scaling properties of instruction tuning for Chinese LLMs.Identifies varying sensitivity of abilities to scaling factors.Investigates impact of data quantity, model size, and data construction.

Instruction Tuning for Large Language Models: A Survey

Aug 21, 2023
SZ
Shengyu Zhang
🏛️ Zhejiang University | Shannon.AI | Nanyang Technological University | Amazon

This study addresses the fundamental alignment gap between large language models’ (LLMs) pretraining objective—next-token prediction—and human-centric instruction-following requirements. It systematically surveys instruction tuning techniques, analyzing methodological evolution, strategies for constructing high-quality instruction-output pairs, multi-stage training paradigms, and cross-modal/domain adaptation pathways. Key determinants of generalization and controllability—such as data diversity, format consistency, and task coverage—are identified. Innovatively, the work introduces the first structured, knowledge-graph-style survey integrating theoretical foundations, practical frameworks, and critical reflection. It explicitly delineates current limitations—including instruction bias and the absence of standardized evaluation metrics—and proposes future research directions: scalable alignment, dynamic instruction synthesis, and causally grounded controllable generation. The resulting synthesis has become a benchmark reference in the LLM alignment community.

Bridging gap between model prediction and user instruction adherenceReviewing methodologies, datasets, and applications of supervised fine-tuningSurveying instruction tuning techniques for large language models

Chain-of-Instructions: Compositional Instruction Tuning on Large Language Models

Feb 18, 2024
SH
S. Hayati
🏛️ University of Minnesota | Amazon | Grammarly

Existing large language models exhibit limited zero-shot generalization to multi-step compositional tasks—such as cross-lingual summarization—due to their reliance on single-step instruction paradigms. To address this, we propose the Chain-of-Instruction (CoI) paradigm, which explicitly models complex tasks as sequential chains of input-output subtasks, thereby enhancing end-to-end reasoning through stepwise decomposition. We formally define CoI for the first time, departing from conventional single-step instruction fine-tuning. Leveraging only existing instruction datasets, we construct CoI training samples and apply standard supervised fine-tuning (SFT), requiring no architectural modifications or reinforcement learning. Extensive experiments demonstrate that CoI-tuning consistently improves zero-shot generalization across compositional tasks—including long-chain reasoning, cross-lingual generation, and multi-hop question answering—with scalable gains. It significantly outperforms strong baselines across multiple benchmarks, establishing a new state-of-the-art in structured instruction learning.

Complex Task HandlingMulti-lingual SummarizationMulti-step Tasks

Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning

Aug 27, 2024
SK
Simran Kaur
🏛️ Princeton University | Meta

To address the high cost, labor-intensive nature, and limited diversity of human-curated supervised fine-tuning (SFT) data, this paper proposes Instruct-SkillMix—a fully automated pipeline that leverages LLM-based metacognitive prompting to disentangle core skills embedded in instructions, followed by stochastic skill recombination to generate high-quality, diverse SFT samples. This work establishes, for the first time, a skill-disentanglement-and-random-recomposition paradigm for SFT data generation. It empirically reveals SFT’s extreme sensitivity to low-quality samples and elucidates the intrinsic cause of crowd-sourced SFT data underperformance. Using only 4K synthetically generated samples to fine-tune LLaMA-3-8B-Base, our method achieves a 42.76% win rate on AlpacaEval 2.0—competitive with state-of-the-art models such as Claude 3 Opus—while incurring a total cost under $600.

Addresses challenges in naive crowd-sourced instruction-tuning datasetsAutomates creation of diverse SFT data for instruction-followingImproves LLM performance on benchmarks with minimal data

A Survey on Data Selection for LLM Instruction Tuning

Feb 04, 2024
JW
Jiahao Wang
🏛️ Harbin Institute of Technology | Chinese Academy of Sciences

This study addresses the challenge of efficiently selecting high-quality data subsets for instruction tuning to enhance LLM performance while reducing training costs. We systematically survey mainstream instruction datasets and propose, for the first time, a taxonomy of data selection methodologies specifically designed for LLM instruction tuning—establishing a “quality-driven” paradigm to supplant the conventional “quantity-driven” approach. Our framework categorizes strategies into four classes: model-based feedback, uncertainty estimation, diversity optimization, and instruction complexity modeling. We further design a comprehensive downstream evaluation suite—including AlpacaEval and MT-Bench—to enable consistent, multi-dimensional assessment. Through unified benchmarking of over 30 selection methods, we empirically demonstrate that retaining only 10–30% of high-quality samples achieves performance comparable to full-dataset tuning, significantly mitigating critical issues such as evaluation inconsistency.

Enhancing LLM instruction tuning effectiveness through data selectionImproving LLM instruction-following capabilities via optimized data subsetsReducing training costs by selecting high-quality instruction datasets

Latest Papers

What's happening recently
View more

Existing instruction datasets, though reaching millions in scale, suffer from insufficient coverage across task types, limited diversity across knowledge domains, and inadequate depth in instruction complexity—constraining fine-tuned models’ generalization to complex instructions and low-resource domains. To address this, we propose a “coverage–depth” co-enhancement paradigm, introducing a closed-loop data construction framework integrating hierarchical annotation, informative seed selection, evolutionary synthesis, and defect-driven targeted generation. This framework shifts emphasis from mere quantity to qualitative advancement, substantially expanding the information-theoretic boundary of instruction distributions. Leveraging it, we curate a high-quality dataset of 1.5 million instructions. Empirical evaluation across multiple foundation models and benchmarks (e.g., MT-Bench, AlpacaEval) demonstrates systematic improvements in instruction-following capability—particularly on challenging tasks requiring long-chain reasoning and cross-domain inference.

Enhancing instruction dataset coverage and depthImproving complex instruction-following in rare domainsSystematic framework for high-quality instruction data construction

Existing instruction-tuning datasets often conflate world knowledge acquired during pretraining with the instruction-following capabilities developed during post-training, thereby limiting fine-tuning effectiveness. This work proposes CoDIT (Contrastive Decoding for Instruction Tuning), a method that leverages contrastive decoding between a post-trained model and its pretrained counterpart to suppress shared world knowledge and amplify pure instruction-following behavior. CoDIT achieves, for the first time, the distillation of a “chat vector” from parameter space into textual space, effectively disentangling and transferring instruction-following ability in a manner compatible across diverse model architectures. Models trained on datasets constructed via CoDIT substantially outperform those trained on directly generated data or existing public instruction-tuning benchmarks, demonstrating significantly enhanced instruction-following performance.

instruction tuninginstruction-following capabilitieslarge language models

This work addresses the challenge of adapting large language models to specialized domains, which is often constrained by the scarcity of high-quality, low-cost domain-specific instruction-tuning data. The authors propose a zero-shot instruction synthesis framework that, for the first time, integrates Bloom’s cognitive taxonomy with task-aware keywords to automatically generate diverse, multi-domain instructions. To ensure the professionalism and reliability of the synthesized data, the framework incorporates a self-consistency verification mechanism. Notably, the approach requires no human annotation and successfully produces high-quality instruction data across seven specialized domains. Models fine-tuned with this synthetic data significantly outperform those trained using existing data synthesis methods.

data synthesisdomain-specificinstruction tuning

When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following

Sep 25, 2025
KH
Keno Harada
🏛️ The University of Tokyo | Kyoto University | University of the Ryukyus

Prior work lacks systematic evaluation of large language models’ (LLMs) ability to simultaneously follow multiple instructions—a critical yet underexplored capability. Method: We introduce two dedicated benchmarks—ManyIFEval for text generation and StyleMBPP for code generation—covering diverse multi-instruction combinations. To enable efficient assessment, we propose lightweight regression models (e.g., logistic regression) that predict model performance using features such as instruction count, generalizing to unseen instruction sets and arbitrary instruction numbers. Results: Experiments reveal a pronounced performance degradation with increasing instruction count; our models achieve accurate predictions (within ~10% error) using only 300–500 samples, drastically reducing evaluation overhead. Our core contributions are: (1) the first systematic benchmarking framework for multi-instruction following, (2) a generalizable performance prediction methodology, and (3) the first quantitative characterization of the inverse relationship between instruction count and adherence performance.

Developing models to estimate performance on unseen instruction combinationsEvaluating LLMs' ability to follow multiple instructions simultaneouslyMeasuring performance degradation as instruction count increases systematically

From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding

Jun 04, 2025
CZ
Chiwei Zhu
🏛️ University of Science and Technology of China | Metastone Technology

To address the scarcity, low diversity, and weak real-world grounding of high-quality instruction data in large language model (LLM) alignment, this paper proposes the *attributed grounding* framework, which synergistically integrates top-down user-context attribution with bottom-up, web-document-driven joint context–instruction generation. Our method employs instruction provenance analysis, multi-granularity context modeling, web retrieval, and structured prompt engineering to construct an end-to-end synthetic pipeline—enabling, for the first time, scalable generation of cognitively inspired yet empirically grounded complex instructions. We release SynthQuestions, a million-scale, high-quality instruction dataset. Evaluated on multiple alignment benchmarks, models trained on SynthQuestions achieve significant performance gains, with improvements consistently scaling with the volume of underlying web corpora.

Enhancing instruction complexity with attributed grounding methodsGenerating diverse synthetic instructions for LLM alignmentOvercoming limited grounding sources in instruction synthesis

Hot Scholars

JJ

Jiaya Jia

Chair Professor, HKUST; Adjunct Prof., CUHK
Artificial IntelligenceComputer VisionDeep Learning
HW

Hongning Wang

Associate Professor, Department of Computer Science and Technology, Tsinghua University
Machine LearningInformation RetrievalLarge Language Models
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
TS

Tat-Seng Chua

National University of Singapore
Multimedia Information RetrievalLive Social Media Analysis
NM

Niklas Muennighoff

Stanford University
large language modelsartificial intelligencemachine learning