Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the diminishing returns of serial computation during mid-training of large language models, where sequential training struggles to yield sustained improvements in downstream performance. To overcome this scaling bottleneck, we propose a multi-branch parallel training framework that allocates the computational budget across independent branches. Guided by validation metrics, optimal checkpoints are selected and fused via intra-layer and inter-layer parameter averaging. Grounded in the principle of compatible diversity, this approach leverages cross-trajectory averaging to effectively cancel residual errors inherent in single-sequence training. Experiments demonstrate that, under matched training budgets, our method significantly enhances aggregate downstream performance. Moreover, these gains scale consistently with increased budgets and transfer effectively into post-training stages.
📝 Abstract
Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since additional serial compute yields little further downstream improvement and can even degrade some capabilities, which places a practical ceiling on how much compute mid-training absorbs. We revisit how this compute should be allocated to a single run or multiple similar optimizations. We find that branches forked from a shared checkpoint under various controlled recipe reaches measurably different regions of parameter space, and establish a form of compatible diversity that extending one run cannot supply. Therefore, we introduce Trajectory Soup, which distributes a mid-training budget over several independent branches, and consolidates strongest checkpoints selected on validation through intra- and inter-trajectory averaging into a single model. A local bias and variance analysis separates the two averaging levels, showing that inter-trajectory averaging removes residual error beyond the reach of averaging within a trajectory, while checkpoint selection carries a bias that bounds how many checkpoints are worth merging. Across model scales, learning-rate schedules, token budgets, and trajectory counts, Trajectory Soup improves aggregate downstream performance over the strongest single-trajectory average under matched budgets and keeps improving as budgets expand, with the advantage preserved after an identical post-training pipeline. These results position trajectory allocation and merging as a practical way to extend the compute-scaling frontier of mid-training beyond serial saturation.
Problem

Research questions and friction points this paper is trying to address.

Mid-training
Compute scaling
Large language models
Trajectory diversity
Serial saturation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trajectory Soup
Mid-training
Compute-scaling
Model Averaging
Diverse Trajectories
🔎 Similar Papers
No similar papers found.