On the Impact of Fine-Tuning on Chain-of-Thought Reasoning

📅 2024-11-22
🏛️ arXiv.org
📈 Citations: 5
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates how fine-tuning affects the chain-of-thought (CoT) reasoning capabilities of large language models (LLMs), particularly its detrimental impact on reasoning faithfulness and generalization. Through controlled, comparative experiments across four standard CoT benchmarks, we systematically evaluate prevalent fine-tuning paradigms—including supervised fine-tuning (SFT), RLHF, and Q-LoRA—and quantitatively demonstrate, for the first time, that while fine-tuning improves task accuracy, it consistently degrades CoT faithfulness by a significant margin. To diagnose this phenomenon, we introduce a fine-grained faithfulness evaluation framework and empirically establish that fine-tuning induces systematic shifts in internal reasoning mechanisms. Our core contribution is the first causal identification of fine-tuning as a driver of CoT faithfulness degradation, accompanied by a reproducible diagnostic toolkit—laying foundational groundwork for developing trustworthy, interpretable LLM reasoning systems.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and RobustnessReasoning under Uncertainty: Causality

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Large language models have emerged as powerful tools for general intelligence, showcasing advanced natural language processing capabilities that find applications across diverse domains. Despite their impressive performance, recent studies have highlighted the potential for significant enhancements in LLMs' task-specific performance through fine-tuning strategies like Reinforcement Learning with Human Feedback (RLHF), supervised fine-tuning (SFT), and Quantized Low-Rank Adapters (Q-LoRA) method. However, previous works have shown that while fine-tuning offers significant performance gains, it also leads to challenges such as catastrophic forgetting and privacy and safety risks. To this end, there has been little to no work in extit{understanding the impact of fine-tuning on the reasoning capabilities of LLMs}. Our research investigates the effect of fine-tuning on the reasoning abilities of LLMs, addressing critical questions regarding the impact of task-specific fine-tuning on overall reasoning capabilities, the influence of fine-tuning on Chain-of-Thought (CoT) reasoning performance, and the implications for the faithfulness of CoT reasonings. By exploring these dimensions, our study shows the impact of fine-tuning on LLM reasoning capabilities, where the faithfulness of CoT reasoning, on average across four datasets, decreases, highlighting potential shifts in internal mechanisms of the LLMs resulting from fine-tuning processes.
Problem

Research questions and friction points this paper is trying to address.

Impact of fine-tuning on LLM reasoning capabilities
Effect of fine-tuning on Chain-of-Thought performance
Faithfulness changes in CoT reasoning post-fine-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Investigates fine-tuning impact on LLM reasoning
Analyzes Chain-of-Thought performance changes post-tuning
Quantifies faithfulness decline in CoT reasoning
🔎 Similar Papers
No similar papers found.