Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high computational cost of iterative pre-training interventions and the tendency of post-hoc fine-tuning to induce capability degradation and "reality drift." To overcome these limitations, this work proposes a weight grafting method. The core idea involves training adapters on pre-training checkpoints and transferring them to post-trained models, thereby approximating the effects of genuine pre-training interventions. Technically, the approach integrates synthetic document fine-tuning, weight update grafting, and cross-checkpoint model alignment, enabling adapter reuse without requiring re-post-training. Experimental results demonstrate that this method effectively injects target beliefs while halving both reality drift and capability loss, closely approximating faithful pre-training performance. Consequently, it substantially reduces experimental costs and facilitates rapid iterative development for large-scale models.
πŸ“ Abstract
Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideally, synthetic documents would be mixed into pre- or mid-training, but every change to a pre-training corpus must be followed by a full post-training run before its effect can be measured, making iteration slow and expensive. Common practice instead applies SDF to an already post-trained model. This is known to leave artifacts and degrade capabilities, and, as we show, it makes the model treat fabricated entities unrelated to the documents as real, a failure we call reality drift. We propose grafting: train the SDF adapter on the pre-trained checkpoint, then add the learned weight update to the post-trained model, which approximates the faithful approach while reusing the existing post-training. We demonstrate this by installing false facts, training misaligned model organisms and applying a constitutional mid-training intervention, across model families up to 284B parameters. Grafting installs the target belief as strongly as SDF on the post-trained model while reducing both reality drift and the loss of preference coherence by more than half on average, and it stays closer to a faithful mid-training run. Because grafting requires no post-training, the same adapter can be applied to any later checkpoint, enabling researchers to iterate quickly on pre-training interventions at the cost of a single fine-tuning run.
Problem

Research questions and friction points this paper is trying to address.

pre-training interventions
synthetic document fine-tuning
reality drift
model alignment
post-trained models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Grafting
Synthetic Document Fine-tuning
Reality Drift
Pre-training Interventions
Model Alignment