Score
Design and implement meta-training procedures that learn initial LoRA adapter parameterizations so that only adapter weights are fine-tuned for fast few‑shot adaptation; set up and run first‑order MAML (FOMAML) style optimization over adapter parameters to produce adaptable initializations. Build and evaluate adapter-based few‑shot adaptation workflows, measuring adaptation speed, stability, and cross‑task or cross‑domain generalization after a small number of target updates.
This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.
This study investigates the capacity of small language models to effectively use tools without relying on complex adaptation mechanisms. Focusing on Llama-3.2-3B-Instruct, the authors systematically evaluate four adaptation strategies—hypernetwork-generated LoRA weights, few-shot prompting, document-based prompting, and value-guided beam search—across four tool-use benchmarks. Experimental results demonstrate that few-shot prompting yields a 21.5% performance gain, document prompting contributes an additional 5.0%, while hypernetwork-generated LoRA weights show no significant improvement. Notably, the 3B-parameter model achieves 79.7% of GPT-5’s average performance at only one-tenth of the inference latency. These findings underscore the pivotal role of prompt engineering in enabling efficient tool use with lightweight models and offer a promising direction for resource-constrained settings.
To address the insufficient generalization of pretrained models when fine-tuned on few-shot, high-dimensional binary classification tasks, this paper proposes a novel fine-tuning framework based on weight matrix reparameterization. The method couples low-rank adaptation (LoRA) with a base-model rescaling mechanism and employs random matrix theory to model the generalization behavior of high-dimensional classifiers, thereby revealing how rescaling governs spectral distribution and generalization bounds. Theoretically, the approach substantially mitigates overfitting by controlling the effective rank and condition number of the classifier’s weight matrix. Empirically, it consistently improves performance across multiple binary classification benchmarks and large language model (LLM) fine-tuning tasks, demonstrating particularly pronounced generalization gains under extreme data scarcity.
To address representation collapse and uneven expert load arising from homogeneous Mixture-of-Experts LoRA (MoE-LoRA) in parameter-efficient fine-tuning (PEFT), this paper proposes the Heterogeneous Adapter Mixture (MoA) framework. MoA dynamically integrates structurally diverse lightweight adapters to enable expert specialization. It introduces two novel paradigms—Soft MoA (continuous weighted fusion) and Sparse MoA (top-k expert activation)—thereby overcoming structural homogeneity constraints. By tightly coupling LoRA, MoE, and dynamic gating/weighted fusion mechanisms, MoA achieves task-adaptive sparse expert activation and fine-grained output aggregation. Experiments demonstrate that MoA improves average accuracy by 1.8% across multiple downstream tasks under comparable parameter budgets. Notably, Sparse MoA attains near-zero performance degradation while activating fewer than 15% of experts, substantially enhancing both the efficiency and representational capacity of PEFT.
LoRA adapters often converge to suboptimal solutions near initialization, resulting in poor generalization and low robustness to merging and pruning. To address this, we propose CoTo, a progressive training strategy that introduces stochastic adapter deactivation—dynamically increasing activation probability during training to balance loss landscape exploration and optimization. We theoretically establish that CoTo enhances inter-layer dropout stability and linear mode connectivity. Furthermore, we devise the first Shapley-value-based framework grounded in cooperative game theory to quantify the marginal contribution of each adapter. CoTo is lightweight and fully compatible with mainstream variants such as QLoRA and AdaLoRA. Experiments demonstrate that CoTo improves single-task fine-tuning accuracy by 1.8% on average, boosts multi-task adapter merging accuracy by 4.2%, reduces post-pruning performance degradation by 37%, and cuts training overhead by 15%.
Existing LoRA-MoE studies suffer from inconsistent model architectures, datasets, hyperparameters, and evaluation protocols, hindering fair comparative analysis. Method: We introduce LoRALib, the first systematic benchmark for LoRA-MoE evaluation—integrating 17 model architectures, 680 LoRA modules across 40 downstream tasks, with standardized data formats, fine-tuning pipelines, and hyperparameter configurations; built upon OpenCompass to ensure reproducible large-scale assessment. Contribution/Results: Our key finding is that a task-correlation-aware LoRA selection mechanism significantly improves MoE performance and cross-task generalization. Extensive experiments establish LoRA-MoE as the current state-of-the-art paradigm. All code, data, and trained models are publicly released.
In multi-task and federated learning settings, existing LoRA-based approaches suffer from inefficient parameter sharing across multiple adapters, redundant training of matrix A, and underutilization of matrix B. Method: This paper proposes ALoRA and its federated extension, Fed-ALoRA, which theoretically identify matrix B as the primary carrier of knowledge encoding and transfer, while matrix A captures task-specific characteristics. Accordingly, it introduces an asymmetric sharing architecture: a single shared matrix B paired with multiple task- or client-specific matrices A, further enhanced by heterogeneous rank adaptation to improve generalization under client and task heterogeneity. The method integrates low-rank decomposition, multi-task learning, and federated optimization. Results: Extensive experiments on commonsense reasoning, mathematical reasoning, and multi-task/federated NLP benchmarks demonstrate that ALoRA and Fed-ALoRA achieve accuracy comparable to or surpassing state-of-the-art methods, while yielding more balanced performance across tasks.
This paper addresses the challenge of efficient adaptation of large language models (LLMs). It systematically reviews the architectural evolution and adaptation techniques across Meta AI’s LLaMA series (v1–v4), covering foundational models, Mixture-of-Experts (MoE) variants, and multimodal extensions. To tackle LLM adaptation efficiency, it proposes the first unified, structured comparison of five mainstream Parameter-Efficient Fine-Tuning (PEFT) methods—LoRA, QLoRA, LLaMA-Adapter V1/V2, and LLaMA-Excitor—integrating quantization, low-rank decomposition, and adapter injection strategies across model scales from 7B to 288B parameters. Experiments demonstrate that updating only 0.1%–3% of parameters achieves performance on par with or exceeding full fine-tuning on instruction-following, multimodal understanding, and domain-specific tasks (e.g., healthcare, law). The core contribution is a comprehensive, generation-spanning analytical framework for LLaMA and PEFT, revealing mechanistic insights into high-performance transfer via minimal parameter updates—thereby providing both theoretical grounding and practical paradigms for lightweight LLM deployment.