🤖 AI Summary
To address LoRA’s limited generalization under low-rank constraints and its inferiority to full-parameter fine-tuning, this paper proposes MELoRA—a miniature ensemble-based low-rank adaptation method. MELoRA freezes the pretrained LLM weights and trains multiple ultra-lightweight (rank-1) mini-LoRA adapters in parallel, forming an ensemble. By synergistically combining low-rank matrix decomposition with ensemble learning, it significantly enhances representational capacity while drastically reducing trainable parameters. Theoretical analysis establishes a tighter error upper bound for MELoRA compared to standard LoRA. Empirical evaluation across diverse tasks shows that MELoRA outperforms LoRA on natural language understanding with only 1/8 of its parameters, and achieves comparable or superior performance on instruction-following tasks using merely 1/36 of LoRA’s parameters. To our knowledge, MELoRA is the first PEFT framework to integrate ensemble learning into low-rank adaptation architectures, effectively overcoming the generalization bottleneck inherent in single-adapter approaches.
📝 Abstract
Parameter-efficient fine-tuning (PEFT) is a popular method for tailoring pre-trained large language models (LLMs), especially as the models' scale and the diversity of tasks increase. Low-rank adaptation (LoRA) is based on the idea that the adaptation process is intrinsically low-dimensional, i.e., significant model changes can be represented with relatively few parameters. However, decreasing the rank encounters challenges with generalization errors for specific tasks when compared to full-parameter fine-tuning. We present MELoRA, a mini-ensemble low-rank adapters that uses fewer trainable parameters while maintaining a higher rank, thereby offering improved performance potential. The core idea is to freeze original pretrained weights and train a group of mini LoRAs with only a small number of parameters. This can capture a significant degree of diversity among mini LoRAs, thus promoting better generalization ability. We conduct a theoretical analysis and empirical studies on various NLP tasks. Our experimental results show that, compared to LoRA, MELoRA achieves better performance with 8 times fewer trainable parameters on natural language understanding tasks and 36 times fewer trainable parameters on instruction following tasks, which demonstrates the effectiveness of MELoRA.