FusionBench: A Comprehensive Benchmark of Deep Model Fusion

📅 2024-06-05
🏛️ arXiv.org
📈 Citations: 38
✨ Influential: 6
📄 PDF
🤖 AI Summary
Lack of a unified, reproducible evaluation benchmark hinders rigorous validation of effectiveness and robustness in deep model fusion. To address this, we introduce FusionBench—the first comprehensive benchmark specifically designed for deep model fusion—covering 26 tasks across open-vocabulary image classification, text classification, and generation. It integrates 74 fine-tuned models and 16 fusion methods. We formally define, for the first time, a systematic fusion evaluation paradigm comprising three categories: prediction-level ensembling, parameter-level merging, and component-level hybridization, accompanied by standardized task suites, model configurations, and evaluation protocols. Implemented in PyTorch, FusionBench supports both LoRA and full-parameter fine-tuning and incorporates mainstream algorithms including Task Arithmetic, DARE, SLERP, and Pareto Merging. Empirical analysis reveals significant differences in robustness across fusion methods under distribution shift. The codebase and documentation are publicly available and have been widely adopted as the de facto standard in the field.

Technology Category

Machine Learning: Mixture of Experts (MoE)Intelligent Robots: Multimodal Perception & Sensor FusionNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP Models

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Deep model fusion is an emerging technique that unifies the predictions or parameters of several deep neural networks into a single model in a cost-effective and data-efficient manner. This enables the unified model to take advantage of the original models' strengths, potentially exceeding their performance. Although a variety of deep model fusion techniques have been introduced, their evaluations tend to be inconsistent and often inadequate to validate their effectiveness and robustness against distribution shifts. To address this issue, we introduce FusionBench, which is the first comprehensive benchmark dedicated to deep model fusion. FusionBench covers a wide range of tasks, including open-vocabulary image classification, text classification, and text-to-text generation. Each category includes up to eight tasks with corresponding task-specific models, featuring both full fine-tuning and LoRA fine-tuning, as well as models of different sizes, to ensure fair and balanced comparisons of various multi-task model fusion techniques across different tasks, model scales, and fine-tuning strategies. We implement and evaluate a broad spectrum of deep model fusion techniques. These techniques range from model ensemble methods, which combine the predictions to improve the overall performance, to model merging, which integrates different models into a single one, and model mixing methods, which upscale or recombine the components of the original models. FusionBench now contains 26 distinct tasks, 74 fine-tuned models, and 16 fusion techniques, and we are committed to consistently expanding the benchmark with more tasks, models, and fusion techniques. In addition, we offer a well-documented set of resources and guidelines to aid researchers in understanding and replicating the benchmark results. Homepage https://github.com/tanganke/fusion_bench
Problem

Research questions and friction points this paper is trying to address.

Evaluates deep model fusion techniques for consistency and robustness
Compares fusion methods across diverse tasks and model scales
Provides a unified library for implementing and testing new fusion techniques
Innovation

Methods, ideas, or system contributions that make the work stand out.

FusionBench provides a unified library for model fusion
It offers a comprehensive benchmark for evaluating fusion methods
The benchmark includes diverse tasks and model settings
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.