meta-learn lora adapters

Design and implement meta-training procedures that learn initial LoRA adapter parameterizations so that only adapter weights are fine-tuned for fast few‑shot adaptation; set up and run first‑order MAML (FOMAML) style optimization over adapter parameters to produce adaptable initializations. Build and evaluate adapter-based few‑shot adaptation workflows, measuring adaptation speed, stability, and cross‑task or cross‑domain generalization after a small number of target updates.

meta-learnloraadapters

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.

adapter placementdominant adaptation modulegradient energy

This study investigates the capacity of small language models to effectively use tools without relying on complex adaptation mechanisms. Focusing on Llama-3.2-3B-Instruct, the authors systematically evaluate four adaptation strategies—hypernetwork-generated LoRA weights, few-shot prompting, document-based prompting, and value-guided beam search—across four tool-use benchmarks. Experimental results demonstrate that few-shot prompting yields a 21.5% performance gain, document prompting contributes an additional 5.0%, while hypernetwork-generated LoRA weights show no significant improvement. Notably, the 3B-parameter model achieves 79.7% of GPT-5’s average performance at only one-tenth of the inference latency. These findings underscore the pivotal role of prompt engineering in enabling efficient tool use with lightweight models and offer a promising direction for resource-constrained settings.

few-shot adaptationmodel adaptationprompt engineering

$α$-LoRA: Effective Fine-Tuning via Base Model Rescaling

Oct 24, 2025
AE
Aymane El Firdoussi
🏛️ EPFL | TII

To address the insufficient generalization of pretrained models when fine-tuned on few-shot, high-dimensional binary classification tasks, this paper proposes a novel fine-tuning framework based on weight matrix reparameterization. The method couples low-rank adaptation (LoRA) with a base-model rescaling mechanism and employs random matrix theory to model the generalization behavior of high-dimensional classifiers, thereby revealing how rescaling governs spectral distribution and generalization bounds. Theoretically, the approach substantially mitigates overfitting by controlling the effective rank and condition number of the classifier’s weight matrix. Empirically, it consistently improves performance across multiple binary classification benchmarks and large language model (LLM) fine-tuning tasks, demonstrating particularly pronounced generalization gains under extreme data scarcity.

Enhancing reparameterization methods for transfer learningImproving generalization ability of fine-tuned modelsValidating effectiveness through theoretical and experimental approaches

To address representation collapse and uneven expert load arising from homogeneous Mixture-of-Experts LoRA (MoE-LoRA) in parameter-efficient fine-tuning (PEFT), this paper proposes the Heterogeneous Adapter Mixture (MoA) framework. MoA dynamically integrates structurally diverse lightweight adapters to enable expert specialization. It introduces two novel paradigms—Soft MoA (continuous weighted fusion) and Sparse MoA (top-k expert activation)—thereby overcoming structural homogeneity constraints. By tightly coupling LoRA, MoE, and dynamic gating/weighted fusion mechanisms, MoA achieves task-adaptive sparse expert activation and fine-grained output aggregation. Experiments demonstrate that MoA improves average accuracy by 1.8% across multiple downstream tasks under comparable parameter budgets. Notably, Sparse MoA attains near-zero performance degradation while activating fewer than 15% of experts, substantially enhancing both the efficiency and representational capacity of PEFT.

Addresses representation collapse in homogeneous MoE-LoRA architecturesEnhances knowledge transfer to downstream tasks via heterogeneous adaptersSolves expert load imbalance in parameter-efficient fine-tuning

Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation

Jun 06, 2025
ZZ
Zhuang Zhan
🏛️ Southern University of Science and Technology | City University of Hong Kong | Zhejiang University | University of Surrey | Hong Kong University of Science and Technology | New York University

LoRA adapters often converge to suboptimal solutions near initialization, resulting in poor generalization and low robustness to merging and pruning. To address this, we propose CoTo, a progressive training strategy that introduces stochastic adapter deactivation—dynamically increasing activation probability during training to balance loss landscape exploration and optimization. We theoretically establish that CoTo enhances inter-layer dropout stability and linear mode connectivity. Furthermore, we devise the first Shapley-value-based framework grounded in cooperative game theory to quantify the marginal contribution of each adapter. CoTo is lightweight and fully compatible with mainstream variants such as QLoRA and AdaLoRA. Experiments demonstrate that CoTo improves single-task fine-tuning accuracy by 1.8% on average, boosts multi-task adapter merging accuracy by 4.2%, reduces post-pruning performance degradation by 37%, and cuts training overhead by 15%.

Enhances model generalization and downstream operationsImproves Low-rank adaptation (LoRA) to avoid suboptimal minimaProposes progressive training strategy for balanced optimization

Latest Papers

What's happening recently
View more

LoRALib: A Standardized Benchmark for Evaluating LoRA-MoE Methods

Sep 14, 2025
SW
Shaoheng Wang
🏛️ Zhejiang University of Technology | The City University of New York

Existing LoRA-MoE studies suffer from inconsistent model architectures, datasets, hyperparameters, and evaluation protocols, hindering fair comparative analysis. Method: We introduce LoRALib, the first systematic benchmark for LoRA-MoE evaluation—integrating 17 model architectures, 680 LoRA modules across 40 downstream tasks, with standardized data formats, fine-tuning pipelines, and hyperparameter configurations; built upon OpenCompass to ensure reproducible large-scale assessment. Contribution/Results: Our key finding is that a task-correlation-aware LoRA selection mechanism significantly improves MoE performance and cross-task generalization. Extensive experiments establish LoRA-MoE as the current state-of-the-art paradigm. All code, data, and trained models are publicly released.

Addressing insufficient cross-task generalization in LoRA fine-tuningEnabling fair comparison across models, datasets, and hyperparametersStandardizing evaluation of LoRA-MoE methods lacking unified benchmarks

Rethinking Parameter Sharing for LLM Fine-Tuning with Multiple LoRAs

Sep 29, 2025
HB
Hao Ban
🏛️ University at Buffalo

In multi-task and federated learning settings, existing LoRA-based approaches suffer from inefficient parameter sharing across multiple adapters, redundant training of matrix A, and underutilization of matrix B. Method: This paper proposes ALoRA and its federated extension, Fed-ALoRA, which theoretically identify matrix B as the primary carrier of knowledge encoding and transfer, while matrix A captures task-specific characteristics. Accordingly, it introduces an asymmetric sharing architecture: a single shared matrix B paired with multiple task- or client-specific matrices A, further enhanced by heterogeneous rank adaptation to improve generalization under client and task heterogeneity. The method integrates low-rank decomposition, multi-task learning, and federated optimization. Results: Extensive experiments on commonsense reasoning, mathematical reasoning, and multi-task/federated NLP benchmarks demonstrate that ALoRA and Fed-ALoRA achieve accuracy comparable to or surpassing state-of-the-art methods, while yielding more balanced performance across tasks.

Enabling federated fine-tuning with shared B matricesProposing asymmetric LoRA design for multi-task adaptationReevaluating parameter sharing in multi-LoRA fine-tuning

Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey

Oct 14, 2025
AA
Abdulhady Abas Abdullah
🏛️ University of Kurdistan Hewler | Queen Mary University | Torrens University Australia | University of Technology Sydney | University of Kurdistan Sanandaj | TU Dortmund University | Tehran University

This paper addresses the challenge of efficient adaptation of large language models (LLMs). It systematically reviews the architectural evolution and adaptation techniques across Meta AI’s LLaMA series (v1–v4), covering foundational models, Mixture-of-Experts (MoE) variants, and multimodal extensions. To tackle LLM adaptation efficiency, it proposes the first unified, structured comparison of five mainstream Parameter-Efficient Fine-Tuning (PEFT) methods—LoRA, QLoRA, LLaMA-Adapter V1/V2, and LLaMA-Excitor—integrating quantization, low-rank decomposition, and adapter injection strategies across model scales from 7B to 288B parameters. Experiments demonstrate that updating only 0.1%–3% of parameters achieves performance on par with or exceeding full fine-tuning on instruction-following, multimodal understanding, and domain-specific tasks (e.g., healthcare, law). The core contribution is a comprehensive, generation-spanning analytical framework for LLaMA and PEFT, revealing mechanistic insights into high-performance transfer via minimal parameter updates—thereby providing both theoretical grounding and practical paradigms for lightweight LLM deployment.

Analyzing architectures, performance, and real-world applicationsReviewing parameter-efficient fine-tuning methods for LLMsSurveying the evolution of Meta's LLaMA model series

Hot Scholars

YL

Yong Luo

Wuhan University
Artifical IntelligenceMachine LearningData MiningPattern Classification and Search
AK

Alexander Koller

Professor of Computational Linguistics, Saarland University, Saarland Informatics Campus
Computational linguisticsartificial intelligence
XQ

Xinchi Qiu

Meta, University of Cambridge
GenAIPrivacy-preserving MLAI RobustnessML Systems