Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining
This study addresses whether layer intervention responses during language model pretraining reflect persistent organizational structures or transient training effects. To investigate this, the work employs single-block identity bypasses and fixed teacher-forced contexts to conduct a longitudinal comparative analysis across multiple checkpoint trajectories. The findings reveal that the depth-wise ranking of layer sensitivity remains persistently stable throughout training, whereas intervention magnitudes are dynamically redistributed. Furthermore, the balancing mechanism between local missing updates and downstream responses varies across models. These results demonstrate that longitudinal layer sensitivity is structured yet non-static, thereby challenging the generalizability of intervention-based conclusions drawn from single checkpoints.