Don't Half-listen: Capturing Key-part Information in Continual Instruction Tuning

📅 2024-03-15
🏛️ arXiv.org
📈 Citations: 5
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address catastrophic forgetting (CF) and instruction overfitting in continual instruction tuning (CIT) of large language models (LLMs)—which degrade task generalization and instruction-following capability—this paper proposes a task-aware dynamic tuning framework. Methodologically, it (1) identifies semantically decisive instruction fragments via Key Part Information Gain (KPIG), enabling fine-grained task awareness; (2) introduces a dynamic data replay and target refinement mechanism that preserves task essence rather than superficial patterns; and (3) establishes the first dual-metric evaluation system—P-score (measuring generalization) and V-score (assessing instruction adherence). Experiments demonstrate that our approach significantly outperforms existing baselines on both seen and unseen tasks, effectively mitigating CF and overfitting while improving instruction-following accuracy and cross-task generalization performance.

Technology Category

Machine Learning: Life-Long and Continual LearningSearch and Optimization: Learning to SearchNatural Language Processing: Safety and Robustness

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
Instruction tuning for large language models (LLMs) can drive them to produce results consistent with human goals in specific downstream tasks. However, the process of continual instruction tuning (CIT) for LLMs may bring about the catastrophic forgetting (CF) problem, where previously learned abilities are degraded. Recent methods try to alleviate the CF problem by modifying models or replaying data, which may only remember the surface-level pattern of instructions and get confused on held-out tasks. In this paper, we propose a novel continual instruction tuning method based on Key-part Information Gain (KPIG). Our method computes the information gain on masked parts to dynamically replay data and refine the training objective, which enables LLMs to capture task-aware information relevant to the correct response and alleviate overfitting to general descriptions in instructions. In addition, we propose two metrics, P-score and V-score, to measure the generalization and instruction-following abilities of LLMs. Experiments demonstrate our method achieves superior performance on both seen and held-out tasks.
Problem

Research questions and friction points this paper is trying to address.

Address catastrophic forgetting in continual instruction tuning
Improve task-aware information capture in LLMs
Measure generalization and instruction-following abilities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses Key-part Information Gain for dynamic data replay
Refines training objective to capture task-aware information
Introduces P-score and V-score for ability measurement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Meituan | Institute of Information Engineering, Chinese Academy of Sciences | National University of Defense Technology
Yongquan He
Yongquan He
Meituan, China
X
Xuancheng Huang
Meituan, China
M
Minghao Tang
Institute of Information Engineering, Chinese Academy of Sciences, China
L
Lingxun Meng
Meituan, China
X
Xiang Li
Meituan, China
W
Wei Lin
Meituan, China
W
Wenyuan Zhang
Institute of Information Engineering, Chinese Academy of Sciences, China
Y
Yifu Gao
National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology