🤖 AI Summary
To address representation collapse and uneven expert load arising from homogeneous Mixture-of-Experts LoRA (MoE-LoRA) in parameter-efficient fine-tuning (PEFT), this paper proposes the Heterogeneous Adapter Mixture (MoA) framework. MoA dynamically integrates structurally diverse lightweight adapters to enable expert specialization. It introduces two novel paradigms—Soft MoA (continuous weighted fusion) and Sparse MoA (top-k expert activation)—thereby overcoming structural homogeneity constraints. By tightly coupling LoRA, MoE, and dynamic gating/weighted fusion mechanisms, MoA achieves task-adaptive sparse expert activation and fine-grained output aggregation. Experiments demonstrate that MoA improves average accuracy by 1.8% across multiple downstream tasks under comparable parameter budgets. Notably, Sparse MoA attains near-zero performance degradation while activating fewer than 15% of experts, substantially enhancing both the efficiency and representational capacity of PEFT.
📝 Abstract
Recent studies integrate Low-Rank Adaptation (LoRA) and Mixture-of-Experts (MoE) to further enhance the performance of parameter-efficient fine-tuning (PEFT) methods in Large Language Model (LLM) applications. Existing methods employ emph{homogeneous} MoE-LoRA architectures composed of LoRA experts with either similar or identical structures and capacities. However, these approaches often suffer from representation collapse and expert load imbalance, which negatively impact the potential of LLMs. To address these challenges, we propose a emph{heterogeneous} extbf{Mixture-of-Adapters (MoA)} approach. This method dynamically integrates PEFT adapter experts with diverse structures, leveraging their complementary representational capabilities to foster expert specialization, thereby enhancing the effective transfer of pre-trained knowledge to downstream tasks. MoA supports two variants: extbf{(i)} extit{Soft MoA} achieves fine-grained integration by performing a weighted fusion of all expert outputs; extbf{(ii)} extit{Sparse MoA} activates adapter experts sparsely based on their contribution, achieving this with negligible performance degradation. Experimental results demonstrate that heterogeneous MoA outperforms homogeneous MoE-LoRA methods in both performance and parameter efficiency. Our project is available at https://github.com/DCDmllm/MoA.