instruction tuning

Designs, builds, and evaluates instruction artifacts and pipelines — including prompt/instruction templates, instruction corpora for fine-tuning, instruction-following behaviors, instruction parsing and semantics, instruction selection and scheduling algorithms, and assembly- or instruction-set-level layouts — to steer and optimize the behavior of a computational system. Work includes authoring and editing instructions, defining instruction semantics and interpretation, engineering instruction-following and instruction-understanding mechanisms, and measuring how instruction choices (from high-level prompts to ISA decisions) affect system outputs and performance.

instructiontuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.8
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Instruction Tuning for Large Language Models: A Survey

Aug 21, 2023
SZ
Shengyu Zhang
🏛️ Zhejiang University | Shannon.AI | Nanyang Technological University | Amazon

This study addresses the fundamental alignment gap between large language models’ (LLMs) pretraining objective—next-token prediction—and human-centric instruction-following requirements. It systematically surveys instruction tuning techniques, analyzing methodological evolution, strategies for constructing high-quality instruction-output pairs, multi-stage training paradigms, and cross-modal/domain adaptation pathways. Key determinants of generalization and controllability—such as data diversity, format consistency, and task coverage—are identified. Innovatively, the work introduces the first structured, knowledge-graph-style survey integrating theoretical foundations, practical frameworks, and critical reflection. It explicitly delineates current limitations—including instruction bias and the absence of standardized evaluation metrics—and proposes future research directions: scalable alignment, dynamic instruction synthesis, and causally grounded controllable generation. The resulting synthesis has become a benchmark reference in the LLM alignment community.

Bridging gap between model prediction and user instruction adherenceReviewing methodologies, datasets, and applications of supervised fine-tuningSurveying instruction tuning techniques for large language models

Must-Read Papers

Most classic and influential ideas
View more

Boosting Instruction Following at Scale

Oct 16, 2025
BE
Ben Elder
🏛️ IBM T.J. Watson Research

Large language models (LLMs) suffer from a pronounced decline in instruction-following accuracy as the number of concurrent instructions increases—a critical limitation for complex, multi-step prompting. To address this, we propose Instruction Boosting: an interpretable, post-hoc reweighting mechanism grounded in generative outputs. Our method introduces a quantified conflict scoring model that, for the first time, identifies semantic tension among instructions as the primary cause of performance degradation and provides actionable, diagnostic feedback on instruction conflicts. To rigorously evaluate multi-instruction adherence, we construct SCALEDIF, a high-scale benchmark comprising instruction combinations ranging from 2 to 10. Experiments demonstrate that Instruction Boosting improves instruction-following accuracy by +7.0 percentage points in dual-instruction settings and +4.0 points in ten-instruction settings, substantially mitigating multi-instruction performance decay. This work advances prompt engineering with both theoretical insight—revealing instruction conflict as a fundamental bottleneck—and a practical, deployable tool for robust multi-instruction execution.

Addressing performance degradation with increasing instruction volumeImproving LLM instruction following reliability through post-generation methodsQuantifying conflict between multiple instructions to explain performance trends

Improving Instruction-Following in Language Models through Activation Steering

Oct 15, 2024
AS
Alessandro Stolfo
🏛️ ETH Zürich | Microsoft Research

This work addresses the limited capability of large language models (LLMs) to adhere to fine-grained instruction constraints—such as formatting, length, and keyword requirements—and their poor generalization across zero-shot or cross-model settings. To this end, we propose activation steering: a lightweight, inference-time intervention that computes layer-wise neural activation differences between instruction-present and instruction-absent conditions, yielding interpretable, transferable, and composable instruction vectors. Crucially, no model fine-tuning is required. Our key contribution is the first formulation of instructions as cross-model-transferable activation-difference vectors, enabling vector composition (e.g.,叠加 multiple constraints) and foundation-model enhancement. Extensive evaluation across four mainstream LLMs demonstrates substantial improvements in instruction-following accuracy. The method supports constraint-aware generation without explicit instructions, concurrent multi-constraint control, and knowledge transfer from instruction-tuned models to base models.

Controlling output format, length, and word inclusion constraintsEnhancing instruction-following in language models via activation steeringTransferring steering vectors from tuned to base models

This study investigates the impact of instruction tuning on large language models’ performance across two programming interaction paradigms: code completion (“Flow”) and instruction-to-code generation (“Command”). Through human error categorization, generative fidelity metrics, and analysis of intermediate fine-tuning checkpoints, the work systematically demonstrates for the first time that while instruction tuning substantially enhances instruction-following capabilities, it concurrently degrades code infilling performance—a trade-off the authors term the “instruction tuning tax.” The paper articulates seven key findings and four practical implications, underscoring that instruction tuning is not a “free lunch” and offering novel insights for the design of AI-powered programming tools and model optimization strategies.

code generationinfilling performanceinstruction tuning

Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?

Mar 11, 2024
EZ
Egor Zverev
🏛️ ISTA | Microsoft | CISPA Helmholtz Center for Information Security

Large language models (LLMs) suffer from fundamental security vulnerabilities in safety-critical applications due to the blurred logical boundary between instructions and data, rendering them susceptible to adversarial attacks and lacking rigorous, quantifiable safety evaluation metrics. Method: This work formally defines the “instruction-data separation” capability, introduces a computable empirical metric, and proposes SEP—a dedicated benchmark integrating behavioral analysis, formal semantic modeling, and multi-model comparative experiments. Contribution/Results: We demonstrate that all mainstream LLMs exhibit severe deficiencies in instruction-data separation; conventional prompt engineering and fine-tuning fail to enhance security without compromising utility. This study bridges a critical theoretical gap in LLM security assessment by establishing the first reproducible, scalable, and formally grounded safety measurement framework for trustworthy AI.

Instruction-Data DistinctionLanguage Model SecurityVulnerability Quantification

Dynamics of Instruction Tuning: Each Ability of Large Language Models Has Its Own Growth Pace

Oct 30, 2023
CS
Chiyu Song
🏛️ Zhejiang University | Westlake University | Westlake Institute for Advanced Study

This work investigates the mechanisms by which instruction tuning enhances general intelligence in Chinese large language models (LLMs), focusing on how data scale, model size (7B–33B), and data construction methodology (human-authored vs. synthetic) differentially affect multidimensional capabilities—including creative writing, code generation, and logical reasoning. Method: Leveraging a 40k+ multi-capability-annotated instruction dataset, we conduct cross-domain ablation studies to isolate these factors. Contribution/Results: We first reveal that underlying capabilities evolve at independent learning paces; human-authored data remains consistently effective, whereas synthetic data exhibits a performance ceiling; and instruction data demonstrates strong cross-capability generalization. Based on these findings, we propose a quantifiable, efficiency-oriented data construction guideline. Evaluated on two public benchmarks, our approach yields significant performance gains, providing empirical evidence and methodological foundations for capability-targeted LLM optimization.

Explores scaling properties of instruction tuning for Chinese LLMs.Identifies varying sensitivity of abilities to scaling factors.Investigates impact of data quantity, model size, and data construction.

Latest Papers

What's happening recently
View more

The mechanisms by which system prompts influence instruction-tuned models in code generation remain poorly understood, particularly regarding their performance across varying model scales, programming languages, and prompting strategies. This study conducts a large-scale, multi-variable controlled experiment evaluating 360 configurations spanning four instruction-tuned models, five categories of system prompts, three prompting strategies, two programming languages, and multiple temperature settings. The findings reveal that the effectiveness of system prompts is non-monotonic and highly configuration-dependent: notably, few-shot examples can degrade performance in larger models, challenging the conventional wisdom that few-shot prompting consistently outperforms zero-shot. Additionally, Java is found to be more sensitive to prompt design than Python, suggesting the need for language-specific prompting strategies. This work provides empirical foundations and practical guidance for effective prompt engineering in code generation.

code generationcode language modelsinstruction-tuned models

Existing benchmarks often conflate instruction following with task success, hindering accurate assessment of large language models’ true compliance capabilities under complex instructions. This work proposes MOSAIC, a modular framework that, for the first time, decomposes instruction compliance into independently analyzable dimensions. By dynamically synthesizing datasets incorporating up to 20 application-oriented constraints, MOSAIC enables fine-grained, disentangled evaluation. Systematic ablation studies, combined with analyses of constraint composition and positional sensitivity across five mainstream models, reveal non-uniform response patterns dependent on constraint type, count, and placement. The study identifies primacy and recency biases alongside model-specific vulnerabilities, offering critical diagnostic insights to guide the development of more reliable language models.

benchmarkevaluationinstruction compliance

This study investigates whether small-scale language models adhere to user instructions when those instructions conflict with their task capabilities—such as selecting incorrect answers or generating opposite sentiment—and reveals a decoupling between task proficiency and instruction following. To this end, the authors propose a cross-task evaluation paradigm for conflicting instructions and introduce the Instruction Following Failure Rate (IFFR) metric. Systematic experiments on the Qwen model series demonstrate that while smaller models retain task accuracy, they consistently disregard conflicting instructions, whereas larger models exhibit significantly stronger instruction-following behavior. This work provides the first quantitative evidence that task capability does not equate to controllable behavior, offering a novel perspective and methodology for evaluating model controllability.

instruction followinginstruction-conflicting behaviormodel evaluation

This study investigates how large language models internally represent and process instructions during supervised fine-tuning (SFT) and direct preference optimization (DPO). Through causal mediation analysis, we find that instruction representations are highly localized in early network layers and introduce the concept of an “instruction vector”—a representation that effectively guides later layers to select task-relevant information pathways even under conditions of linear non-separability. Our work challenges the prevailing assumption in mechanistic interpretability that internal representations are linearly encoded, and instead proposes a novel method for identifying causal information pathways without relying on linearity. This reveals the instruction vector’s critical role as a selector of task-specific circuits within the model.

instruction representationlanguage modelsmechanistic interpretability

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
JL

Junyang Lin

Qwen Team, Alibaba Group & Peking University
Natural Language ProcessingCross-Modal Representation LearningPretraining
JR

Ji-Rong Wen

Gaoling School of Artificial Intelligence, Renmin University of China
Large Language ModelWeb SearchInformation RetrievalMachine Learning
MY

Min Yang

Bytedance
Vision Language ModelComputer VisionVideo Understanding