Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether large language models with substantially different parameter scales can be effectively fused through training-free weighted averaging. To this end, the authors propose a lightweight dimension-adaptation strategy—aligning parameter spaces via deterministic expansion or truncation—combined with a controllable interpolation mechanism to enable zero-shot fusion of heterogeneous models. Experiments on the Qwen model series and multiple benchmarks demonstrate that even a small interpolation ratio can surpass strong baselines, confirming the method’s efficacy; however, near-balanced interpolation often leads to performance collapse, revealing a seesaw effect in capability transfer. This work establishes a simple yet effective strong baseline for heterogeneous large model fusion and elucidates both its potential and inherent limitations.
📝 Abstract
Can large language models with substantially different parameter spaces be merged by direct weighted averaging, without training or semantic alignment? Existing heterogeneous fusion methods typically introduce distillation, adapters, learned latent spaces, routing, or feature alignment, leaving open whether a simpler recipe can work for genuinely different billion-parameter checkpoints. We revisit this counterintuitive question through training-free dimensional adaptation followed by ratio-controlled interpolation. In union-style merging, we expand the smaller model into the larger parameter space; in intersection-style merging, we truncate the larger model into the smaller parameter space. Across Qwen-family model pairs and benchmarks covering mathematical reasoning, code generation, language understanding, commonsense reasoning, knowledge, and instruction following, deterministic expansion largely preserves the source model function, and small-ratio interpolation can improve over strong source checkpoints by transferring complementary capabilities. However, near-balanced interpolation often collapses, and task-level results reveal a seesaw effect in which gains on some capabilities coexist with regressions on others. These results show that simple parameter averaging, when paired with lightweight dimensional adaptation and carefully controlled ratios, is a surprisingly strong baseline for heterogeneous LLM merging, suggesting that the limits of direct weighted fusion may also bound what more complex heterogeneous merging methods can achieve at scale.
Problem

Research questions and friction points this paper is trying to address.

heterogeneous LLM merging
weighted model averaging
parameter space alignment
training-free fusion
model interpolation
Innovation

Methods, ideas, or system contributions that make the work stand out.

heterogeneous LLM merging
weight averaging
dimensional adaptation
parameter interpolation
training-free fusion
🔎 Similar Papers
J
Jiahe Fan
University of Science and Technology of China
Y
Yinghao Hou
University of Science and Technology of China
S
Si Chen
School of Information Science and Technology, Department of Automation, University of Science and Technology of China
A
Aiyuan Zhang
University of Science and Technology of China
Hong Xie
Hong Xie
University of Science and Technology of China (USTC)
Data Science/MiningOnline Learning
D
Defu Lian
School of Computer Science and Technology, University of Science and Technology of China