task-private lora

Designs and implements per-task low-rank adapter modules (LoRA-style) that are attached to a shared base model to isolate task-specific parameter subspaces and reduce semantic entanglement across tasks. Builds and analyzes methods for parameter-efficient task addition, adapter management, and mitigation of cross-task interference.

task-privatelora

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts

Jul 08, 2025
SY
Samin Yeasar Arnob
🏛️ McGill University | Mila | Microsoft | Université de Montréal | ServiceNow | Georgia Institute of Technology

This work addresses the challenges of multi-task adaptation and merging in parameter-efficient fine-tuning. We propose a modular architecture based on sparse adapters, integrating low-rank decomposition with structured weight sparsity to enable task-specific adapter training with minimal parameter updates—eliminating the need for downstream fine-tuning and supporting plug-and-play multi-task merging. Our key contribution is a streamlined, more efficient sparse training mechanism, both conceptually and empirically superior to LoRA and full-parameter fine-tuning. Experiments demonstrate the first successful joint merging of adapters across 20 diverse NLP tasks, yielding substantial gains in in-distribution performance. The approach consistently outperforms LoRA and full-model merging baselines in multi-task merging scenarios. However, cross-task generalization remains limited and warrants further investigation.

Comparing sparse adapters with LoRA and full fine-tuningExploring sparse adapters for scalable expert mergingInvestigating merging properties across multiple NLP tasks

FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts

Oct 09, 2025
HZ
Heming Zou
🏛️ Tsinghua University | Tianjin University

LoRA suffers from parameter interference during efficient fine-tuning, while MoE-enhanced LoRA mitigates intra-task correlations but incurs routing overhead and fails to address inter-task interference in multi-task merging. To tackle this, we propose FlyLoRA—an implicit mixture-of-experts (MoE) framework inspired by the Drosophila olfactory circuit. FlyLoRA replaces explicit routing with a frozen, sparse random projection matrix that jointly realizes rank-level expert activation and down-projection. This design preserves ultra-low parameter count while significantly enhancing task disentanglement and cross-task generalization. Extensive experiments across four domains—general knowledge, scientific question answering, mathematical reasoning, and code generation—demonstrate that FlyLoRA consistently outperforms state-of-the-art LoRA variants and MoE-based adaptations. Results validate its superior parameter efficiency, robustness against parameter interference, and strong generalization capability across diverse downstream tasks.

Addresses parameter interference in LoRA fine-tuningImproves task decoupling without explicit router parametersMitigates intra-task and inter-task interference simultaneously

This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.

adapter placementdominant adaptation modulegradient energy

ThanoRA: Task Heterogeneity-Aware Multi-Task Low-Rank Adaptation

May 24, 2025
JL
Jian Liang
🏛️ Wuhan University

To address task heterogeneity loss, subspace interference and collapse, and inference overhead induced by parameter merging in multi-task LoRA, this paper proposes the first low-rank adaptation framework that jointly preserves task heterogeneity and subspace independence. Our method introduces: (1) a task-aware LoRA subspace initialization mechanism that explicitly models inter-task differences; (2) an orthogonality regularization constraint on LoRA subspaces to suppress cross-task interference; and (3) a joint training paradigm featuring full parameter sharing, routing-free design, and mergeable adapters. Crucially, the approach incurs zero inference overhead and introduces no additional parameters. Extensive experiments demonstrate that it consistently outperforms strong baselines—including MoE-LoRA—across both multimodal and pure-text multi-task benchmarks, achieving significant gains in generalization performance and task synergy efficiency.

Addresses task heterogeneity with specialized LoRA subspacesEnables multi-task adaptation preserving LoRA inference efficiencyMitigates subspace interference during multi-task training

This work addresses the limitation of standard LoRA, which employs a uniform rank across all layers despite their varying contributions to task adaptation. The authors propose a lightweight, gradient-variance-based rank allocation strategy: prior to fine-tuning, they perform eight calibration backward passes and use the gradient variance of LoRA-B matrices as a proxy for layer informativeness to proportionally distribute a fixed rank budget. This approach introduces no additional parameters, training overhead, or architectural modifications, while yielding interpretable rank assignments. Furthermore, by employing a simplified diagonal approximation of the empirical Fisher information matrix (eFIM) computed solely over LoRA adapters, memory consumption is reduced by approximately 256×. Experiments demonstrate performance on par with standard LoRA on GLUE (88.6 vs. 88.7) and LLaMA-3-8B commonsense reasoning tasks (68.5 vs. 68.7), confirming its efficacy.

layer informativenessLow-rank adaptationparameter efficiency

Latest Papers

What's happening recently
View more

This work addresses the parameter inefficiency of conventional LoRA-based continual learning, where each task is assigned a separate adapter despite significant low-rank redundancy across tasks. The authors propose LiteLoRA, a novel approach that leverages a plug-in gating mechanism and subspace overlap analysis to selectively reuse existing adapters or instantiate new ones in a dynamic, task-aware manner. By identifying and exploiting shared low-rank subspaces among tasks, LiteLoRA achieves competitive or superior performance compared to state-of-the-art methods on standard continual learning benchmarks, while reducing the number of active adapters by 20% to 70%, thereby substantially improving parameter efficiency.

Adapter RedundancyContinual LearningFine-Tuning

This study addresses the challenges of weight interference and retraining costs in merging multiple LoRA adapters by proposing the READ method. Built upon fixed factorization and a unidirectional coupling mechanism, READ enforces balanced canonical form constraints that enable new skills to read legacy inputs without overwriting existing outputs. This design facilitates zero-overhead skill composition during inference without requiring retraining. Experimental results demonstrate that READ surpasses baseline methods by over 20 points on benchmarks such as SuperGLUE while effectively preserving model singularity. By enabling seamless multi-skill integration within parameter-efficient fine-tuning frameworks, this work establishes a novel paradigm for adapter-based model composition.

interferenceLoRA compositionLow-rank adapters

This work addresses the significant performance degradation often observed after merging Low-Rank Adaptation (LoRA) modules, a problem exacerbated by existing approaches that can only assess merge compatibility post-training, leading to wasted computational resources. The paper introduces MergeProbe, the first method capable of predicting post-merge performance retention early in training. MergeProbe leverages lightweight analysis of the alignment between low-rank updates and original gradients, as well as their perturbation of shared representations. It supports diverse composition strategies—including merging, reweighting, pruning, and routing—and consistently outperforms existing interference-aware merging techniques on the MERGE-PEFT benchmark across five domains: mathematics, code, science, instruction following, and safety. Notably, MergeProbe achieves superior average and worst-case performance retention while incurring substantially lower deployment overhead than full task routing.

adapter interferenceLoRAmergeability

Existing methods struggle to effectively merge multiple task-specific LoRA adapters into a single low-rank adapter without causing capability fragmentation or violating the low-rank structure. This work proposes a novel “Compress-then-Merge” (CtM) paradigm: it first constructs a shared r-dimensional subspace from the LoRA weights and orthogonally projects each adapter onto this subspace, then performs standard merging within the resulting r×r core coordinate space, followed by truncated SVD to strictly enforce the target rank constraint. Evaluated across multiple models and tasks, CtM significantly outperforms existing single-LoRA baselines and substantially narrows the performance gap with full-parameter merging, achieving for the first time an efficient and rank-preserving LoRA fusion.

adapter compressionLoRAlow-rank adaptation