model merging

Designs, builds, and analyzes algorithms that combine multiple trained models’ parameters, energies, or predictions into a single or compact representation using techniques such as weight-space interpolation, model averaging/soups, probabilistic and energy-based combinations (e.g., product-of-experts, POE), and optimized merge coefficients. This work includes selecting or trimming influential weights, reducing the dimensionality of the merge, performing coefficient optimization without original training data, and producing unified weights or merged predictors to integrate task-specific knowledge or reduce effects like catastrophic forgetting.

modelmerging

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of effectively merging multiple single-task models into a unified multitask model without relying on additional fine-tuning data, while rigorously evaluating the statistical validity of task-wise update directions. The authors formulate model fusion as a probabilistic inference problem in parameter space and, for the first time, cast individual single-task models as energy-based expert models within a Product-of-Experts (PoE) framework. By identifying a critical mismatch between the implicit Gaussian assumptions in existing methods and the empirically observed heavy-tailed residual distributions, they propose a heavy-tailed PoE based on the Cauchy distribution. This approach consistently outperforms current fusion strategies across diverse tasks and architectures, demonstrating that heavy-tailed modeling is essential for enhancing fusion performance.

energy-based modelsheavy-tailed residualsmodel merging

Why Do More Experts Fail? A Theoretical Analysis of Model Merging

May 27, 2025
ZW
Zijing Wang
🏛️ Northeastern University | LMU Munich

This work addresses the scalability bottleneck in model merging—specifically, the performance degradation observed as the number of experts increases. We establish, for the first time, a theoretical framework grounded in Gaussian width and approximate kinematics, revealing parameter-space saturation as the fundamental limiting factor; we further prove that performance gains exhibit strictly concave decay and admit a unique optimal merging threshold. Building on this insight, we propose Reparameterized Heavy-Tailed (RHT) merging, which alleviates saturation constraints via heavy-tailed reparameterization of expert weights. Extensive evaluation across 12 knowledge-intensive and general-purpose benchmarks demonstrates that RHT significantly delays performance decay and raises the upper bound for multi-task fusion. The implementation is open-sourced. To our knowledge, this is the first theoretically grounded paradigm for scalable model merging, offering provable guarantees on convergence behavior and capacity limits.

Introduces method to enhance merged model performanceInvestigates scalability limits of merging multiple expert modelsProves upper bound on model merging due to parameter constraints

This work addresses the challenge of efficiently fusing independently trained neural network models for capability reuse without access to original training data and with minimal optimization. The authors propose a novel paradigm of direct weight-space fusion: for single-task settings, they introduce the reference-free C²M³ alignment algorithm; for multi-task scenarios, they develop a framework comprising TSV low-rank decomposition, MASS input-adaptive routing, and MERGE³ evolutionary fusion, grounded in gradient-based approximations of task vectors. By innovatively integrating Frank-Wolfe optimization, item response theory for evaluation, and task vector analysis, the method substantially mitigates task interference and computational overhead—reducing evaluation costs by up to 50×—while maintaining strong performance, all without requiring any original training data.

model mergingmulti-task learningsingle-task setting

FREE-Merging: Fourier Transform for Model Merging with Lightweight Experts

Nov 25, 2024
SZ
Shenghe Zheng
🏛️ Harbin Institute of Technology

To address the challenge of task interference in model merging—where performance degradation and deployment overhead hinder simultaneous optimization—this work identifies, for the first time, that interference manifests prominently in the frequency domain, whereas existing methods operate solely in the spatial domain and thus suffer from limited efficacy. We propose a lightweight, Fourier-transform-based expert-augmented fusion framework: (1) a novel frequency-domain filtering mechanism to suppress harmful fine-tuning signals; (2) dynamically activated low-rank expert modules that compensate for information loss at zero training cost; and (3) a unified cross-modal fusion architecture. Evaluated across CV, NLP, and multimodal benchmarks, our method consistently outperforms state-of-the-art approaches, achieving a 37% inference speedup, 52% reduction in parameter storage, and preserving ≥98.6% single-task accuracy.

Addresses task interference in model merging via frequency domain analysisIntroduces FREE-Merging framework to balance cost, latency, and performanceProposes FR-Merging to filter harmful frequency interference efficiently

MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs

Feb 03, 2025
YZ
Yuhang Zhou
🏛️ University of Maryland | Amazon

To address the challenges of integrating multi-domain heterogeneous expert large language models—namely, architectural incompatibility, severe parameter interference, and high fine-tuning costs—this paper proposes a unified Mixture-of-Experts (MoE) model merging framework. Our key contributions are threefold: (1) a novel heterogeneous expert alignment and mapping mechanism enabling seamless integration of both homogeneous and heterogeneous experts; (2) a parameter-interference-resilient weighted fusion strategy coupled with a lightweight dynamic routing heuristic, drastically reducing reliance on task-specific fine-tuning; and (3) multi-objective performance distillation to jointly optimize domain specialization and general-purpose capability. Evaluated on diverse benchmarks—including mathematical reasoning and code generation—our method outperforms existing state-of-the-art merging approaches, reduces fine-tuning cost by over 60%, and achieves substantial gains in generalization and robustness.

Expert ModelsModel FusionPerformance Optimization

Latest Papers

What's happening recently
View more

This work addresses the lack of a unified theoretical foundation in existing model merging approaches and the opacity of hyperparameters in open-source fine-tuned models, which together hinder the predictability of merged model performance. Leveraging L2-stability theory, the study establishes the first unified generalization framework to systematically analyze the generalization capability of merged heterogeneous expert models and proposes actionable fine-tuning strategies to enhance mergeability. Through parameter-space merging, derivation of generalization bounds, and large-scale vision experiments on ResNet and ViT architectures, the authors validate the critical influence of hyperparameters on merging performance across 20 and 8 tasks, respectively. Theoretical predictions align closely with empirical results, significantly improving the predictability and effectiveness of model merging.

fine-tuned modelsgeneralizationheterogeneous hyperparameters

This work addresses the lack of systematic understanding in efficiently fusing large language models (LLMs) fine-tuned with lightweight adapters in multi-task learning, particularly regarding the trade-offs among ensembling, merging, and routing strategies. The study systematically evaluates three parameter-efficient fusion approaches—output ensembling, parameter averaging, and input-dependent routing—and demonstrates that non-uniform fusion consistently outperforms uniform methods, with routing yielding significant performance gains despite its higher computational cost. To reconcile this efficiency–performance trade-off, the authors propose a low-overhead expert selection mechanism that combines clustering with greedy subset selection, achieving near-optimal performance while substantially reducing computational overhead, thereby striking an effective balance between model efficacy and efficiency.

ensemblingmergingmulti-task learning

This work proposes a modular expert recombination framework to address the limitations of existing model fusion approaches, which typically treat task-specific models as monolithic entities and lack component-level granularity and module reusability. The framework constructs a reusable library of component-level experts and employs a lightweight dynamic routing network to adaptively assemble an optimal sub-model at inference time based on the input. The fusion process is formulated as a bi-objective optimization problem, and a surrogate-assisted evolutionary algorithm efficiently searches for Pareto-optimal configurations. Extensive experiments demonstrate that the proposed method consistently outperforms strong baselines across diverse model scales, task types, and fine-tuning strategies, achieving superior generalization, inference efficiency, and storage economy.

fine-grained mergingheterogeneous tasksmergeability

Model Merging via Multi-Teacher Knowledge Distillation

Dec 24, 2025
SA
Seyed Arshan Dalili
🏛️ The Pennsylvania State University

Model merging, as a lightweight multi-task learning paradigm, lacks theoretical grounding and robust optimization mechanisms for generalization under data-scarce, label-free, and heterogeneous task-distribution settings. To address this, we propose a novel model merging framework based on multi-teacher knowledge distillation, jointly optimizing a student model on scarce unlabeled data. We establish the first flatness-aware PAC-Bayes generalization bound tailored for model merging, introducing “cross-task heterogeneity” to quantify prior-target distribution mismatch. Furthermore, we formulate coefficient scaling as an optimizable KL-divergence minimization problem and integrate Sharpness-Aware Minimization (SAM) to enhance training stability. Our method achieves state-of-the-art performance on multi-task vision and NLP benchmarks, demonstrating significantly improved generalization and robustness under distribution shift. The implementation is publicly available.

Develops method to find flat minima for robust merged model performanceEstablishes generalization bounds for model merging across heterogeneous data distributionsFrames model merging as multi-teacher distillation on scarce unlabeled data

This work challenges the conventional practice of merging models at the point of optimal validation loss by systematically investigating how the training duration of expert models affects the performance of merged large language models. Through multi-stage training checkpoints across five domains and three model scales, evaluated with five merging methods, the study reveals a strong dependence between merging efficacy and training length. It finds that simple averaging suffers significant degradation during overfitting, whereas sparse merging methods achieve peak performance well beyond the validation-optimal step. Drawing a theoretical analogy to random forests via bias-variance decomposition, the paper proposes jointly optimizing training duration and merging strategy, establishing a new paradigm for efficient model merging.

expert modelslarge language modelsmodel merging

Hot Scholars

ER

Emanuele Rodolà

Professor of Computer Science, Sapienza University of Rome
Machine LearningAudioGeometric Deep LearningGeometry Processing
DC

Donato Crisostomi

Ellis Ph.D. student, Sapienza University of Rome & University of Cambridge
Deep LearningModel MergingRepresentation Alignment
HY

Hung-yi Lee

National Taiwan University
deep learningspoken language understandingspeech processing
SC

Simone Calderara

University of Modena and Reggio Emilia
Machine learningcontinual learningtrackingpattern recognition
HY

Hongxia Yang

Professor, HK Polytechnic University
Machine LearningGenerative AICognitive IntelligenceStatistical Modeling