Finetuning with Sampling: SFT Learns Better Than You Think

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of supervised fine-tuning (SFT), which suffers from weak generalization and catastrophic forgetting due to its reliance on off-policy data. To overcome these bottlenecks, this work proposes a Markov Chain Monte Carlo (MCMC)-based sampling algorithm that formulates sampling as a model-native operator for the first time. Guided by a reference model, the approach progressively transforms off-policy expert data into an on-policy distribution. By reshaping the data distribution rather than modifying the objective function, it enables SFT to effectively leverage privileged information. Experiments on scientific skill acquisition and mathematical reasoning tasks demonstrate that the proposed method allows SFT to achieve performance comparable to reinforcement learning, while significantly mitigating catastrophic forgetting and enhancing out-of-distribution generalization capabilities.
📝 Abstract
Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.
Problem

Research questions and friction points this paper is trying to address.

Supervised Finetuning
Catastrophic Forgetting
Off-policy Data
Post-training
Generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Supervised Finetuning (SFT)
Markov Chain Monte Carlo (MCMC)
Off-policy to On-policy Sampling
Post-training
Catastrophic Forgetting
🔎 Similar Papers
No similar papers found.