From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the long-standing confounding among validation data volume, checkpoint selection criteria, and final-checkpoint performance in supervised fine-tuning (SFT). By modeling checkpoint selection as a decision problem under limited information and fixing training trajectories, this work decouples these three factors to independently quantify their respective gains. It proposes a selection strategy based on generation accuracy and checkpoint consistency, incorporating bootstrapping for confidence interval analysis. Empirical results demonstrate that expanding the validation budget significantly improves test accuracy, and that generative selection criteria outperform negative log-likelihood (NLL). However, the advantage of such criteria over simply using the final checkpoint remains uncertain. Overall, this research provides a rigorous decoupled analytical framework for SFT evaluation.
📝 Abstract
Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fixed-budget comparisons do not by themselves distinguish three empirical claims: whether more validation data improve checkpoint selection, whether a selection rule outperforms validation-loss selection, and whether it improves over simply retaining the final checkpoint. We therefore treat checkpoint selection as a finite-information decision problem. Holding completed training trajectories, candidate checkpoints, and independent test items fixed, we vary the validation budget and separately measure improvement from additional validation data, gain over negative log-likelihood (NLL) selection, and gain over the final checkpoint. Across 60 mathematical SFT trajectories and 19 configurations, increasing the validation budget from 32 to 305-313 examples raises independent-test accuracy by 0.32 percentage points (pp) for generated-accuracy selection and 0.29 pp for checkpoint agreement, with 95% configuration-bootstrap CIs of [0.10, 0.56] and [0.11, 0.50], respectively. At the full validation budget, the two generation-based rules outperform matched NLL selection by 0.71 and 0.85 pp, respectively, while their gains over the final checkpoint remain unresolved. A cross-domain replication on 12 newly trained Commonsense trajectories shows the same qualitative separation: increasing the validation budget from 32 to 1,024 questions improves generated-accuracy and checkpoint-agreement selection by 0.87 and 0.27 pp, while gains over the final checkpoint again remain unresolved. Together, these results show that benefiting from more validation data, outperforming NLL selection, and outperforming the final checkpoint are distinct empirical claims that require separate evidence.
Problem

Research questions and friction points this paper is trying to address.

supervised fine-tuning
checkpoint selection
validation budget
model evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Checkpoint Selection
Supervised Fine-Tuning
Finite-Information Decision Problem
Validation Budget
Generation-Based Selection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yupeng Chang
School of Artificial Intelligence, Jilin University
Wenxuan Zhang
Wenxuan Zhang
Singapore University of Technology and Design
Natural Language ProcessingLarge Language ModelsMultilingual NLP
Y
Yuan Wu
School of Artificial Intelligence, Jilin University; Key Laboratory of Symbolic Computation and Knowledge Engineering, Jilin University