Learning Style, Forgetting Semantics: A Case Study of SFT and RFT on Classification Tasks

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the poorly understood mechanism by which supervised fine-tuning (SFT) induces more severe semantic forgetting than reinforcement fine-tuning (RFT). By comparing SFT and RFT on classification tasks, this work decomposes parameter updates using a linear softmax strategy, combining exact gradient derivations with sequential task simulations to analyze how style drift affects semantic retention. Theoretically, it reveals that SFT causes off-axis drift due to teacher style preferences, leading to forgetting, whereas RFT preserves symmetry to maintain semantics. It rigorously establishes theoretical bounds on their forgetting disparity, proving that under specific conditions, SFT exhibits a strictly positive lower bound on forgetting while RFT achieves zero semantic error. Simulation experiments comprehensively validate these theoretical predictions.
📝 Abstract
Why does supervised fine-tuning (SFT) lead to more forgetting than reinforcement fine-tuning (RFT), even when all teacher demonstrations are semantically correct? We study this question on classification tasks where tokens within each semantic class express the same semantic answer in different styles. The tasks share an underlying semantic rule but differ in their prompt distributions and teachers' stylistic preferences. Using a tractable linear-softmax policy, we derive an exact decomposition of the updates into semantic and style components. We show that, at a common policy and prompt, SFT and RFT have parallel semantic updates but differ in their style dynamics. Starting from a policy with no within-class style preference, RFT with exact policy gradients preserves this symmetry, whereas SFT with a nonuniform teacher develops off-axis style drift along a nonzero task mean under population updates. We use this drift to establish a separation under explicit conditions: for population updates from a common perfectly fitted checkpoint, SFT forgetting admits a strictly positive lower bound over a finite training interval, while RFT retains zero semantic error. Simulations over task sequences support these theoretical predictions.
Problem

Research questions and friction points this paper is trying to address.

Supervised Fine-Tuning
Reinforcement Fine-Tuning
Catastrophic Forgetting
Classification Tasks
Style Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Supervised Fine-Tuning
Reinforcement Fine-Tuning
Catastrophic Forgetting
Style Drift
Linear-Softmax Policy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.