🤖 AI Summary
This study addresses the inability of checkpoint merging to accurately reconstruct the terminal state of sequence training by proposing a second-order precision merging algorithm. Methodologically, the approach achieves high-fidelity model weight fusion through convex combination, gradient descent trajectory analysis, and moment condition matching techniques. Theoretically, it establishes information-theoretic limits and derives a unique optimal coefficient rule. Experimental results demonstrate that the proposed method significantly outperforms the WSM baseline under long context windows. Beyond establishing theoretical reconstruction bounds, this work provides a principled criterion for solving merging coefficients that balances optimal error rates with practical applicability.
📝 Abstract
Saved checkpoints record states along a training trajectory, but generally do not determine the updates at states that would be visited under a different schedule. We study how accurately these checkpoints can reconstruct the endpoint of a sequential reference with prescribed update strengths. Under a common local transition model, two checkpoint-index moment conditions characterize all convex merges that agree with this reference through second order. We then prove an information limit that for nondegenerate profiles, no algorithm using only a fixed-length gradient-descent (GD) history with step size $h$ can achieve $o(h^3)$ endpoint error uniformly over a fixed class of smooth, strongly convex losses. The lower bound follows from two losses with identical GD checkpoint histories but sequential reference endpoints separated by $\Omega(h^3)$. \textbf{Quadratic-Accurate Merging} (QAM) achieves a matching uniform $O(h^3)$ endpoint error bound. Its explicit coefficients also define the unique profile-dependent merge that exactly matches the sequential GD reference across all fixed quadratic objectives. Across two public Adam checkpoint trajectories (SmolLM3-3B and OpenEuroLLM-Prelude-9B), three windows and three profiles per model, and 15 tasks, QAM shows mixed results for short windows and broader advantages over \textbf{Warmup-Stable and Merge} (WSM) for longer windows. Matched-moment GSM8K diagnostics further show that local consistency alone does not fully determine downstream scores. These results characterize the reconstruction limits of saved histories, provide a coefficient rule that attains the optimal rate, and assess its practical utility.