🤖 AI Summary
This study addresses the limitation that transferring capabilities from expert models to general-purpose language models typically relies on training or alignment. To overcome this, we propose a training-free heterogeneous model merging method that eliminates the need for gradient updates and semantic alignment. By leveraging parameter projection and interpolation, our approach directly facilitates cross-role knowledge transfer at the parameter level. Furthermore, we design two core merging strategies: Intersection-Merge and Activate-Prune-Merge. Experimental results demonstrate that the proposed method significantly enhances the performance of general-purpose models across embedding, reranking, and code generation tasks. This work establishes a novel paradigm for the efficient integration of heterogeneous models.
📝 Abstract
Specialized models encode task-oriented behavior, but transferring that behavior to a general language model usually requires training, distillation, or representation alignment. We study whether such ability can instead be transferred directly at the parameter level. We apply two existing training-free heterogeneous merging methods, previously shown to transfer knowledge between general language models, to specialist-to-general transfer, projecting a specialist donor into the recipient's shape and interpolating backbone parameters without gradient updates or semantic alignment. Intersection-Merge (IM) injects a prefix-aligned donor slice matching the recipient shape, while Activate-Prune-Merge (APM) uses forward-pass activation statistics to select which donor dimensions to retain before injection. Across embedding, reranking, reward modeling, and MoE code-specialist transfer, both methods improve the general recipient, showing that simple heterogeneous merging can move capabilities across diverse specialist roles.