🤖 AI Summary
This study addresses the performance degradation of frozen layers during fine-tuning, where updates to unfrozen layers shift input distributions. To mitigate this, we propose Guarded Freezing, a strategy that reveals how network connectivity influences freezing efficacy. We introduce a bimodal scoring mechanism based on capacity and drift, which dynamically applies removal value or drift value to guide weight freezing according to path isolation or residual stream sharing states, thereby overcoming the limitations of conventional static freezing. Extensive experiments across VGG, Qwen, and DINOv3 architectures demonstrate that our approach significantly improves accuracy retention on prior tasks, consistently outperforming baseline methods such as DEFT, Wanda, and Fisher.
📝 Abstract
When adapting pre-trained models through fine-tuning, freezing weights alone might not preserve performance, as updates elsewhere can change the inputs to the frozen core, ultimately affecting overall performance. We first analyse the case where a selected frozen core can be isolated and propose removal-value, a capacity-based score that approximates HOPE's removal cost averaged over removal orders. We show that in VGG-8, cutting paths from trainable neurons into a frozen core makes selection using this score useful: $70\%$ frozen preserves $5.22\pm0.51$ percentage points more old-task accuracy than DEFT at similar new-task accuracy. In transformers, shared residual streams leave paths into frozen neurons open. For this case, we derive drift-value, a forward-only proxy for the output disturbance from updating each weight entry under a local update model. In language models, at 40 epochs, this policy exceeds adapted Wanda and RIA freezing scores in settings with substantial retention loss, while its differences from Fisher remain unresolved. After 160 epochs on Qwen2.5-1.5B, it retains $0.0433\pm0.0102$ more than static Fisher. In DINOv3 vision-transformer adaptation to point clouds, drift-value retains $0.440$ image accuracy versus $0.187$ for a random mask of the same count. These results motivate Guarded Freezing: select by removal-value when incoming paths are cut, and by drift-value when they remain.