Score
Designs, curates, and synthesizes datasets of instruction-style examples—covering instruction–input–output triples, instruction–code pairs, and instruction-linked knowledge items—including construction and maintenance of instruction knowledge bases and task-specific corpora. Defines instruction templates, parsing and encoding schemes, generates synthetic instructions and multi-stage prompt sequences, prepares data for instruction tuning/fine-tuning, and builds evaluation protocols and benchmarks to measure instruction-following behavior.
To address the critical misalignment between large language models (LLMs) and human intent, safety constraints, and domain-specific requirements, this paper proposes an alignment-centric instruction tuning paradigm. Methodologically, it systematically integrates three core components: (1) data construction—encompassing expert annotation, model distillation, and self-improvement; (2) efficient fine-tuning—including full-parameter tuning, LoRA, and prefix tuning; and (3) multidimensional evaluation—featuring automated generation, adaptive optimization, and robustness validation. It innovatively classifies and unifies data construction strategies and establishes a multilingual, multimodal, domain-specific benchmark covering healthcare, law, finance, and other high-stakes fields. The key contribution is the first reusable technical framework and practical guideline that jointly optimizes alignment depth, training efficiency, and evaluation reliability—demonstrably enhancing LLMs’ safety, reliability, and domain adaptability in complex real-world scenarios.
This work investigates the mechanisms by which instruction tuning enhances general intelligence in Chinese large language models (LLMs), focusing on how data scale, model size (7B–33B), and data construction methodology (human-authored vs. synthetic) differentially affect multidimensional capabilities—including creative writing, code generation, and logical reasoning. Method: Leveraging a 40k+ multi-capability-annotated instruction dataset, we conduct cross-domain ablation studies to isolate these factors. Contribution/Results: We first reveal that underlying capabilities evolve at independent learning paces; human-authored data remains consistently effective, whereas synthetic data exhibits a performance ceiling; and instruction data demonstrates strong cross-capability generalization. Based on these findings, we propose a quantifiable, efficiency-oriented data construction guideline. Evaluated on two public benchmarks, our approach yields significant performance gains, providing empirical evidence and methodological foundations for capability-targeted LLM optimization.
This study addresses the fundamental alignment gap between large language models’ (LLMs) pretraining objective—next-token prediction—and human-centric instruction-following requirements. It systematically surveys instruction tuning techniques, analyzing methodological evolution, strategies for constructing high-quality instruction-output pairs, multi-stage training paradigms, and cross-modal/domain adaptation pathways. Key determinants of generalization and controllability—such as data diversity, format consistency, and task coverage—are identified. Innovatively, the work introduces the first structured, knowledge-graph-style survey integrating theoretical foundations, practical frameworks, and critical reflection. It explicitly delineates current limitations—including instruction bias and the absence of standardized evaluation metrics—and proposes future research directions: scalable alignment, dynamic instruction synthesis, and causally grounded controllable generation. The resulting synthesis has become a benchmark reference in the LLM alignment community.
Existing large language models exhibit limited zero-shot generalization to multi-step compositional tasks—such as cross-lingual summarization—due to their reliance on single-step instruction paradigms. To address this, we propose the Chain-of-Instruction (CoI) paradigm, which explicitly models complex tasks as sequential chains of input-output subtasks, thereby enhancing end-to-end reasoning through stepwise decomposition. We formally define CoI for the first time, departing from conventional single-step instruction fine-tuning. Leveraging only existing instruction datasets, we construct CoI training samples and apply standard supervised fine-tuning (SFT), requiring no architectural modifications or reinforcement learning. Extensive experiments demonstrate that CoI-tuning consistently improves zero-shot generalization across compositional tasks—including long-chain reasoning, cross-lingual generation, and multi-hop question answering—with scalable gains. It significantly outperforms strong baselines across multiple benchmarks, establishing a new state-of-the-art in structured instruction learning.
To address the high cost, labor-intensive nature, and limited diversity of human-curated supervised fine-tuning (SFT) data, this paper proposes Instruct-SkillMix—a fully automated pipeline that leverages LLM-based metacognitive prompting to disentangle core skills embedded in instructions, followed by stochastic skill recombination to generate high-quality, diverse SFT samples. This work establishes, for the first time, a skill-disentanglement-and-random-recomposition paradigm for SFT data generation. It empirically reveals SFT’s extreme sensitivity to low-quality samples and elucidates the intrinsic cause of crowd-sourced SFT data underperformance. Using only 4K synthetically generated samples to fine-tune LLaMA-3-8B-Base, our method achieves a 42.76% win rate on AlpacaEval 2.0—competitive with state-of-the-art models such as Claude 3 Opus—while incurring a total cost under $600.
This study addresses the challenge of efficiently selecting high-quality data subsets for instruction tuning to enhance LLM performance while reducing training costs. We systematically survey mainstream instruction datasets and propose, for the first time, a taxonomy of data selection methodologies specifically designed for LLM instruction tuning—establishing a “quality-driven” paradigm to supplant the conventional “quantity-driven” approach. Our framework categorizes strategies into four classes: model-based feedback, uncertainty estimation, diversity optimization, and instruction complexity modeling. We further design a comprehensive downstream evaluation suite—including AlpacaEval and MT-Bench—to enable consistent, multi-dimensional assessment. Through unified benchmarking of over 30 selection methods, we empirically demonstrate that retaining only 10–30% of high-quality samples achieves performance comparable to full-dataset tuning, significantly mitigating critical issues such as evaluation inconsistency.
Existing instruction datasets, though reaching millions in scale, suffer from insufficient coverage across task types, limited diversity across knowledge domains, and inadequate depth in instruction complexity—constraining fine-tuned models’ generalization to complex instructions and low-resource domains. To address this, we propose a “coverage–depth” co-enhancement paradigm, introducing a closed-loop data construction framework integrating hierarchical annotation, informative seed selection, evolutionary synthesis, and defect-driven targeted generation. This framework shifts emphasis from mere quantity to qualitative advancement, substantially expanding the information-theoretic boundary of instruction distributions. Leveraging it, we curate a high-quality dataset of 1.5 million instructions. Empirical evaluation across multiple foundation models and benchmarks (e.g., MT-Bench, AlpacaEval) demonstrates systematic improvements in instruction-following capability—particularly on challenging tasks requiring long-chain reasoning and cross-domain inference.
Existing instruction-tuning datasets often conflate world knowledge acquired during pretraining with the instruction-following capabilities developed during post-training, thereby limiting fine-tuning effectiveness. This work proposes CoDIT (Contrastive Decoding for Instruction Tuning), a method that leverages contrastive decoding between a post-trained model and its pretrained counterpart to suppress shared world knowledge and amplify pure instruction-following behavior. CoDIT achieves, for the first time, the distillation of a “chat vector” from parameter space into textual space, effectively disentangling and transferring instruction-following ability in a manner compatible across diverse model architectures. Models trained on datasets constructed via CoDIT substantially outperform those trained on directly generated data or existing public instruction-tuning benchmarks, demonstrating significantly enhanced instruction-following performance.
This work addresses the challenge of adapting large language models to specialized domains, which is often constrained by the scarcity of high-quality, low-cost domain-specific instruction-tuning data. The authors propose a zero-shot instruction synthesis framework that, for the first time, integrates Bloom’s cognitive taxonomy with task-aware keywords to automatically generate diverse, multi-domain instructions. To ensure the professionalism and reliability of the synthesized data, the framework incorporates a self-consistency verification mechanism. Notably, the approach requires no human annotation and successfully produces high-quality instruction data across seven specialized domains. Models fine-tuned with this synthetic data significantly outperform those trained using existing data synthesis methods.
Prior work lacks systematic evaluation of large language models’ (LLMs) ability to simultaneously follow multiple instructions—a critical yet underexplored capability. Method: We introduce two dedicated benchmarks—ManyIFEval for text generation and StyleMBPP for code generation—covering diverse multi-instruction combinations. To enable efficient assessment, we propose lightweight regression models (e.g., logistic regression) that predict model performance using features such as instruction count, generalizing to unseen instruction sets and arbitrary instruction numbers. Results: Experiments reveal a pronounced performance degradation with increasing instruction count; our models achieve accurate predictions (within ~10% error) using only 300–500 samples, drastically reducing evaluation overhead. Our core contributions are: (1) the first systematic benchmarking framework for multi-instruction following, (2) a generalizable performance prediction methodology, and (3) the first quantitative characterization of the inverse relationship between instruction count and adherence performance.
To address the scarcity, low diversity, and weak real-world grounding of high-quality instruction data in large language model (LLM) alignment, this paper proposes the *attributed grounding* framework, which synergistically integrates top-down user-context attribution with bottom-up, web-document-driven joint context–instruction generation. Our method employs instruction provenance analysis, multi-granularity context modeling, web retrieval, and structured prompt engineering to construct an end-to-end synthetic pipeline—enabling, for the first time, scalable generation of cognitively inspired yet empirically grounded complex instructions. We release SynthQuestions, a million-scale, high-quality instruction dataset. Evaluated on multiple alignment benchmarks, models trained on SynthQuestions achieve significant performance gains, with improvements consistently scaling with the volume of underlying web corpora.