Learning Dynamics of LLM Finetuning

📅 2024-07-15
🏛️ arXiv.org
📈 Citations: 2
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates learning dynamics in large language model (LLM) fine-tuning, focusing on three critical phenomena: (1) cross-task factual transfer and response repetition hallucinations in instruction tuning; (2) the “squeezing effect”—an anomalous decline in target response probability—under prolonged preference optimization; and (3) the superiority of on-policy over off-policy direct preference optimization (DPO). We propose an influence-function-based stepwise attribution method and construct a response-level influence-path model, providing the first unified explanation for these phenomena. We formally identify and characterize the squeezing effect as arising from excessive contraction of the policy distribution, thereby revealing the intrinsic advantage of on-policy training. Leveraging this insight, we design a novel alignment-optimization strategy. Experiments demonstrate that our framework significantly improves fine-tuning stability and alignment performance, offering both theoretical foundations and practical solutions for efficient, controllable LLM fine-tuning.

Technology Category

Search and Optimization: Learning to SearchMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language Models

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Learning dynamics, which describes how the learning of specific training examples influences the model's predictions on other examples, gives us a powerful tool for understanding the behavior of deep learning systems. We study the learning dynamics of large language models during different types of finetuning, by analyzing the step-wise decomposition of how influence accumulates among different potential responses. Our framework allows a uniform interpretation of many interesting observations about the training of popular algorithms for both instruction tuning and preference tuning. In particular, we propose a hypothetical explanation of why specific types of hallucination are strengthened after finetuning, e.g., the model might use phrases or facts in the response for question B to answer question A, or the model might keep repeating similar simple phrases when generating responses. We also extend our framework and highlight a unique"squeezing effect"to explain a previously observed phenomenon in off-policy direct preference optimization (DPO), where running DPO for too long makes even the desired outputs less likely. This framework also provides insights into where the benefits of on-policy DPO and other variants come from. The analysis not only provides a novel perspective of understanding LLM's finetuning but also inspires a simple, effective method to improve alignment performance.
Problem

Research questions and friction points this paper is trying to address.

Analyzes LLM learning dynamics
Explains finetuning-induced hallucinations
Describes DPO's squeezing effect
Innovation

Methods, ideas, or system contributions that make the work stand out.

Analyzes influence accumulation in LLM
Explains hallucination post-finetuning
Introduces squeezing effect in DPO
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
University of British Columbia | Amii