Score
Designs, builds, and analyzes algorithms that combine multiple trained models’ parameters, energies, or predictions into a single or compact representation using techniques such as weight-space interpolation, model averaging/soups, probabilistic and energy-based combinations (e.g., product-of-experts, POE), and optimized merge coefficients. This work includes selecting or trimming influential weights, reducing the dimensionality of the merge, performing coefficient optimization without original training data, and producing unified weights or merged predictors to integrate task-specific knowledge or reduce effects like catastrophic forgetting.
This work addresses fundamental challenges in model merging—including the absence of a unified taxonomy, terminological inconsistency, incomparable methodologies, and difficulties in multi-task fusion under data-unavailable scenarios. We propose the first three-tiered classification paradigm encompassing weight-space fusion, gradient alignment, and task disentanglement. We establish a cross-method reproducible evaluation benchmark and formally define and distinguish the applicability boundaries of “data-agnostic” versus “data-aware” merging. By unifying the theoretical formulations of over 20 state-of-the-art methods—via spectral analysis, normalization sensitivity diagnosis, and task vector geometric modeling—we identify three root causes of merging failure: directional conflict, scale mismatch, and task entanglement. Our framework provides systematic theoretical foundations and principled design guidelines for efficient, lightweight, and interpretable model fusion.
This work addresses the challenge of effectively merging multiple single-task models into a unified multitask model without relying on additional fine-tuning data, while rigorously evaluating the statistical validity of task-wise update directions. The authors formulate model fusion as a probabilistic inference problem in parameter space and, for the first time, cast individual single-task models as energy-based expert models within a Product-of-Experts (PoE) framework. By identifying a critical mismatch between the implicit Gaussian assumptions in existing methods and the empirically observed heavy-tailed residual distributions, they propose a heavy-tailed PoE based on the Cauchy distribution. This approach consistently outperforms current fusion strategies across diverse tasks and architectures, demonstrating that heavy-tailed modeling is essential for enhancing fusion performance.
This work addresses the scalability bottleneck in model merging—specifically, the performance degradation observed as the number of experts increases. We establish, for the first time, a theoretical framework grounded in Gaussian width and approximate kinematics, revealing parameter-space saturation as the fundamental limiting factor; we further prove that performance gains exhibit strictly concave decay and admit a unique optimal merging threshold. Building on this insight, we propose Reparameterized Heavy-Tailed (RHT) merging, which alleviates saturation constraints via heavy-tailed reparameterization of expert weights. Extensive evaluation across 12 knowledge-intensive and general-purpose benchmarks demonstrates that RHT significantly delays performance decay and raises the upper bound for multi-task fusion. The implementation is open-sourced. To our knowledge, this is the first theoretically grounded paradigm for scalable model merging, offering provable guarantees on convergence behavior and capacity limits.
This work addresses the challenge of efficiently fusing independently trained neural network models for capability reuse without access to original training data and with minimal optimization. The authors propose a novel paradigm of direct weight-space fusion: for single-task settings, they introduce the reference-free C²M³ alignment algorithm; for multi-task scenarios, they develop a framework comprising TSV low-rank decomposition, MASS input-adaptive routing, and MERGE³ evolutionary fusion, grounded in gradient-based approximations of task vectors. By innovatively integrating Frank-Wolfe optimization, item response theory for evaluation, and task vector analysis, the method substantially mitigates task interference and computational overhead—reducing evaluation costs by up to 50×—while maintaining strong performance, all without requiring any original training data.
To address the challenge of task interference in model merging—where performance degradation and deployment overhead hinder simultaneous optimization—this work identifies, for the first time, that interference manifests prominently in the frequency domain, whereas existing methods operate solely in the spatial domain and thus suffer from limited efficacy. We propose a lightweight, Fourier-transform-based expert-augmented fusion framework: (1) a novel frequency-domain filtering mechanism to suppress harmful fine-tuning signals; (2) dynamically activated low-rank expert modules that compensate for information loss at zero training cost; and (3) a unified cross-modal fusion architecture. Evaluated across CV, NLP, and multimodal benchmarks, our method consistently outperforms state-of-the-art approaches, achieving a 37% inference speedup, 52% reduction in parameter storage, and preserving ≥98.6% single-task accuracy.
To address the challenges of integrating multi-domain heterogeneous expert large language models—namely, architectural incompatibility, severe parameter interference, and high fine-tuning costs—this paper proposes a unified Mixture-of-Experts (MoE) model merging framework. Our key contributions are threefold: (1) a novel heterogeneous expert alignment and mapping mechanism enabling seamless integration of both homogeneous and heterogeneous experts; (2) a parameter-interference-resilient weighted fusion strategy coupled with a lightweight dynamic routing heuristic, drastically reducing reliance on task-specific fine-tuning; and (3) multi-objective performance distillation to jointly optimize domain specialization and general-purpose capability. Evaluated on diverse benchmarks—including mathematical reasoning and code generation—our method outperforms existing state-of-the-art merging approaches, reduces fine-tuning cost by over 60%, and achieves substantial gains in generalization and robustness.
This work addresses the lack of a unified theoretical foundation in existing model merging approaches and the opacity of hyperparameters in open-source fine-tuned models, which together hinder the predictability of merged model performance. Leveraging L2-stability theory, the study establishes the first unified generalization framework to systematically analyze the generalization capability of merged heterogeneous expert models and proposes actionable fine-tuning strategies to enhance mergeability. Through parameter-space merging, derivation of generalization bounds, and large-scale vision experiments on ResNet and ViT architectures, the authors validate the critical influence of hyperparameters on merging performance across 20 and 8 tasks, respectively. Theoretical predictions align closely with empirical results, significantly improving the predictability and effectiveness of model merging.
This work addresses the lack of systematic understanding in efficiently fusing large language models (LLMs) fine-tuned with lightweight adapters in multi-task learning, particularly regarding the trade-offs among ensembling, merging, and routing strategies. The study systematically evaluates three parameter-efficient fusion approaches—output ensembling, parameter averaging, and input-dependent routing—and demonstrates that non-uniform fusion consistently outperforms uniform methods, with routing yielding significant performance gains despite its higher computational cost. To reconcile this efficiency–performance trade-off, the authors propose a low-overhead expert selection mechanism that combines clustering with greedy subset selection, achieving near-optimal performance while substantially reducing computational overhead, thereby striking an effective balance between model efficacy and efficiency.
This work proposes a modular expert recombination framework to address the limitations of existing model fusion approaches, which typically treat task-specific models as monolithic entities and lack component-level granularity and module reusability. The framework constructs a reusable library of component-level experts and employs a lightweight dynamic routing network to adaptively assemble an optimal sub-model at inference time based on the input. The fusion process is formulated as a bi-objective optimization problem, and a surrogate-assisted evolutionary algorithm efficiently searches for Pareto-optimal configurations. Extensive experiments demonstrate that the proposed method consistently outperforms strong baselines across diverse model scales, task types, and fine-tuning strategies, achieving superior generalization, inference efficiency, and storage economy.
Model merging, as a lightweight multi-task learning paradigm, lacks theoretical grounding and robust optimization mechanisms for generalization under data-scarce, label-free, and heterogeneous task-distribution settings. To address this, we propose a novel model merging framework based on multi-teacher knowledge distillation, jointly optimizing a student model on scarce unlabeled data. We establish the first flatness-aware PAC-Bayes generalization bound tailored for model merging, introducing “cross-task heterogeneity” to quantify prior-target distribution mismatch. Furthermore, we formulate coefficient scaling as an optimizable KL-divergence minimization problem and integrate Sharpness-Aware Minimization (SAM) to enhance training stability. Our method achieves state-of-the-art performance on multi-task vision and NLP benchmarks, demonstrating significantly improved generalization and robustness under distribution shift. The implementation is publicly available.
This work challenges the conventional practice of merging models at the point of optimal validation loss by systematically investigating how the training duration of expert models affects the performance of merged large language models. Through multi-stage training checkpoints across five domains and three model scales, evaluated with five merging methods, the study reveals a strong dependence between merging efficacy and training length. It finds that simple averaging suffers significant degradation during overfitting, whereas sparse merging methods achieve peak performance well beyond the validation-optimal step. Drawing a theoretical analogy to random forests via bias-variance decomposition, the paper proposes jointly optimizing training duration and merging strategy, establishing a new paradigm for efficient model merging.