Predicting and Repairing Merge Collapse in Large Language Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance collapse problem in large language model merging by proposing a variance-based prediction and adaptive repair framework. The research reveals the counter-predictive nature of sign conflicts during task vector arithmetic, motivating the design of the PRISM operator. By integrating task vector statistics, noise modeling, and inter-layer soft thresholding, this work constructs a data-free mechanism for automatic interference mitigation. The proposed method accurately predicts destructive merging behaviors and effectively restores model performance to baseline levels under collapse scenarios. Ultimately, this framework provides a reliable, data-free solution for the safe merging of large models.
📝 Abstract
Large language models fine-tuned from a shared base can be merged by averaging their task vectors, but some merges collapse far below the base model, and common merge operators give no warning before evaluation. We show that one statistic of the specialists' task vectors both predicts this collapse and calibrates its repair. The power that averaging removes equals the variance of the task vectors across specialists, our measure of interference. Under a working noise model, the disturbance that a merge injects grows with the merge coefficient and with interference, yielding a pre-merge score. In our experiments on twenty-two merge configurations from four model families, only destructive merges exceed a threshold on this score. We find that statistics of sign conflict between specialists, a common target of existing merge operators, are anti-predictive. We then predicted the outcomes of fourteen merges before evaluating them, and twelve predictions were correct, including the destructive outcome of a specialist pair pushed past the threshold by continued pretraining. To address this collapse, we introduce PRISM, an operator that averages the task vectors first and then soft-thresholds each layer at a level set by the layer's interference. Without data or tuning, PRISM keeps all five destructive merges above the threshold within evaluation noise of the base model, where plain averaging falls at least 14.4 points below it or collapses entirely. We apply PRISM only above the threshold and keep the plain average for merges below it, which include all fifteen harmless ones. Code is available at https://github.com/js-lee-AI/PRISM.
Problem

Research questions and friction points this paper is trying to address.

Model Merging
Merge Collapse
Large Language Models
Task Vectors
Interference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model Merging
Task Vectors
Interference Prediction
PRISM
Soft-thresholding
🔎 Similar Papers