LoRALib: A Standardized Benchmark for Evaluating LoRA-MoE Methods

📅 2025-09-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing LoRA-MoE studies suffer from inconsistent model architectures, datasets, hyperparameters, and evaluation protocols, hindering fair comparative analysis. Method: We introduce LoRALib, the first systematic benchmark for LoRA-MoE evaluation—integrating 17 model architectures, 680 LoRA modules across 40 downstream tasks, with standardized data formats, fine-tuning pipelines, and hyperparameter configurations; built upon OpenCompass to ensure reproducible large-scale assessment. Contribution/Results: Our key finding is that a task-correlation-aware LoRA selection mechanism significantly improves MoE performance and cross-task generalization. Extensive experiments establish LoRA-MoE as the current state-of-the-art paradigm. All code, data, and trained models are publicly released.

Technology Category

Machine Learning: Mixture of Experts (MoE)Natural Language Processing: (Large) Language ModelsComputer Vision: Large Vision Models

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
As a parameter efficient fine-tuning (PEFT) method, low-rank adaptation (LoRA) can save significant costs in storage and computing, but its strong adaptability to a single task is often accompanied by insufficient cross-task generalization capabilities. To improve this, existing work combines LoRA with mixture-of-experts (MoE) to enhance the model's adaptability through expert modules and routing mechanisms. However, existing LoRA-MoE methods lack unified standards in models, datasets, hyperparameters, and evaluation methods, making it difficult to conduct fair comparisons between different methods. To this end, we proposed a unified benchmark named LoRALib. Specifically, we standardized datasets from $40$ downstream tasks into a unified format, fine-tuned them using the same hyperparameters and obtained $680$ LoRA modules across $17$ model architectures. Based on this LoRA library, we conduct large-scale experiments on $3$ representative LoRA-MoE methods and different LoRA selection mechanisms using the open-sourced testing tool OpenCompass. Extensive experiments show that LoRAMoE performs best, and that prioritizing LoRAs relevant to the target task can further improve the performance of MoE. We hope these findings will inspire future work. Our datasets and LoRA library are available at https://huggingface.co/datasets/YaoLuzjut/LoRAOcean_dataset and https://huggingface.co/YaoLuzjut/models.
Problem

Research questions and friction points this paper is trying to address.

Standardizing evaluation of LoRA-MoE methods lacking unified benchmarks
Addressing insufficient cross-task generalization in LoRA fine-tuning
Enabling fair comparison across models, datasets, and hyperparameters
Innovation

Methods, ideas, or system contributions that make the work stand out.

Standardized datasets from 40 downstream tasks
Fine-tuned 680 LoRA modules across 17 models
Benchmarked 3 LoRA-MoE methods with OpenCompass
🔎 Similar Papers
No similar papers found.
S
Shaoheng Wang
Zhejiang University of Technology
Y
Yao Lu
Zhejiang University of Technology
Y
Yuqi Li
The City University of New York
Y
Yaxin Gao
Zhejiang University of Technology
J
Jiaqi Nie
Zhejiang University of Technology
S
Shanqing Yu
Zhejiang University of Technology
Yingli Tian
Yingli Tian
Distinguished Professor, EE of The City College and CS of the Graduate Center, CUNY
Computer VisionMachine LearningMedical Imaging Analysis
Qi Xuan
Qi Xuan
Professor, Zhejiang University of Technology
AI SecuritySocial NetworkDeep LearningData Mining