🤖 AI Summary
LoRA achieves efficiency in single-task fine-tuning but struggles with cross-task knowledge transfer and relies heavily on task-specific data in multi-task settings. To address this, we propose MeTA-LoRA, a two-stage low-rank adaptation framework: (1) a task-specific stage, where lightweight adapters are learned from minimal per-task samples; and (2) a meta-adaptation stage, where a shared adapter is updated via gradient aggregation across tasks, enabling implicit knowledge transfer. To our knowledge, this is the first work to systematically integrate multi-task knowledge transfer into the LoRA paradigm. Experiments across multi-task and multilingual benchmarks demonstrate that MeTA-LoRA attains performance on par with—or exceeding—that of full-data LoRA fine-tuning, while using only 1%–5% of the task-specific training data. This substantially improves data efficiency and generalization capability of large language models under constrained-data regimes.
📝 Abstract
Low-Rank Adaptation (LoRA) has emerged as one of the most widely used parameter-efficient fine-tuning (PEFT) methods for adapting large language models (LLMs) to downstream tasks. While highly effective in single-task settings, it struggles to efficiently leverage inter-task knowledge in complex multi-task learning scenarios, often requiring substantial task-specific data to achieve optimal performance. To address this limitation, we introduce MeTA-LoRA, a two-stage optimization framework that significantly improves data efficiency in multi-task adaptation. In the first stage, task-specific LoRA adapters are learned using only a few samples from each involved dataset, enabling rapid adaptation without large-scale supervision. In the second stage, the shared LoRA adapter is updated by aggregating gradients from multiple tasks to promote knowledge transfer across tasks, further reducing data usage by leveraging common patterns. In both multi-task learning and multilingual learning scenarios, our method matches or surpasses the performance of traditional full-data LoRA fine-tuning approaches, while using significantly less task-specific data.