🤖 AI Summary
This study addresses the limitation of existing Transformer model merging methods, which neglect functional architecture and consequently induce attention interference and cross-layer error accumulation. We propose Sequential Local Operator Alignment, a method that estimates component behaviors along execution paths using calibration data to align and factorize merged operators layer by layer. This approach introduces a sequential alignment mechanism to suppress cross-layer error propagation and leverages operator factorization for rank expansion, thereby enhancing multi-task capacity. As a training-free merging paradigm, it consistently outperforms existing baselines across diverse multimodal settings and model scales, from CLIP to large language models (LLMs). By eliminating the need for additional training, the proposed method effectively optimizes the trade-off between accuracy and inference cost.
📝 Abstract
Training-free model merging aims to combine multiple fine-tuned models into a single model without further optimization on labeled data. Yet, in transformers, independently merging individual layers can affect a shared attention computation because the query-key and value-output operators depend on composed matrices, overlooking the functional structure. Moreover, when merging earlier components, downstream components receive different activations than they do in the original model, thus, the merged and original execution paths no longer match. In this paper, we introduce Sequential Local Operator Alignment, a training-free method that merges transformers along the execution path of the partially merged model. Our method uses calibration data to estimate the local behavior of each functional component, aligns operators sequentially under the intermediate activation of the partially merged model, and subsequently factorizes the merged operators back into valid transformer parameters. We empirically show that this sequential step reduces error accumulation across layers. Furthermore, the proposed operator factorization step enables rank expansion, providing a principled mechanism for increasing multi-task capacity. We demonstrate that our approach generalizes across modalities, model scales, and varying numbers of tasks, from CLIP and RoBERTa to billion-parameter LLMs, and further extends naturally to the merging of LoRA-fine-tuned models. The results indicate improvements over strong merging baselines without requiring rank expansion, while optional expansion provides a further accuracy-inference-cost trade-off. Project link: https://akansh12.github.io/SLOA-Merge/