Score
Design and train low‑rank adapter (LoRA) modules and associated dual or iterative LoRA training pipelines that allocate independent adapters to distinct latent factors (e.g., content vs. style or motion) so each adapter encodes a disentangled attribute. Build evaluation and recombination procedures that analyze and combine those adapters to achieve controllable model outputs.
Facing the challenge of inefficient navigation and filtering among over 100,000 LoRA adapters hosted on large-scale platforms, this paper proposes a submodular optimization-based adapter selection framework. It formulates adapter retrieval as a combinatorial optimization problem balancing relevance and diversity. Methodologically, we design a differentiable submodular objective function that jointly incorporates low-rank decomposition features from attention layers and semantic similarity metrics, enabling end-to-end optimization. We evaluate the approach quantitatively (retrieval accuracy, diversity score) and qualitatively (generation quality, stylistic coverage) on text-to-image generation tasks. Experiments demonstrate significant improvements across multiple domains: +12.3% Recall@10 in adapter retrieval and +28.6% LPIPS diversity gain, while preserving generation fidelity. This work establishes a novel paradigm for efficient utilization of large-scale lightweight adapter repositories.
To address the slow convergence and suboptimal performance of LoRA in large-model fine-tuning—caused by random weight initialization—this paper proposes a training-free, data-driven initialization method. The core insight is to formulate LoRA initialization as a domain adaptation problem, where activation distribution alignment between pretraining and fine-tuning stages is explicitly enforced; this yields a closed-form solution for directly estimating adapter weights. Moreover, the method supports rank-adaptive decomposition of the up/down projection matrices, enhancing parameterization flexibility. Crucially, it requires neither additional training nor backpropagation—only forward-pass activation statistics and singular value decomposition. Evaluated across image generation, classification, and understanding tasks, the approach accelerates convergence by 1.8× on average and outperforms standard LoRA and existing data-driven initialization methods by +2.3% in average task performance.
Existing methods struggle to efficiently compose an arbitrary number of heterogeneous LoRA adapters in open-world scenarios. This paper proposes LoRAtorio—a training-free, dynamic multi-LoRA composition framework for visual concept combination and controllable generation in text-to-image diffusion models. Its core contributions are: (1) a spatially aware weight matrix derived from latent-space noise prediction discrepancies; (2) a block-wise cosine similarity–driven weighted aggregation strategy; and (3) an extended classifier-free guidance mechanism to mitigate domain shift and enable dynamic adapter selection. Experiments demonstrate that LoRAtorio achieves state-of-the-art performance: +1.3% ClipScore improvement and a 72.43% win rate in GPT-4V pairwise evaluation, while maintaining compatibility across diverse latent diffusion architectures.
This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.
To address performance degradation and concept forgetting caused by rapid switching among multiple LoRA modules, this paper proposes Sparse High-Rank Adapters (SHiRA)—a parameter-efficient, zero-inference-overhead adaptation paradigm. SHiRA fine-tunes only 1–2% of the base model’s parameters, enabling millisecond-scale adapter switching via structured sparse weight updates and fusion-aware training, while supporting collaborative multi-adapter fusion. Theoretical analysis elucidates how high sparsity enhances multi-task synergy. SHiRA is fully compatible with mainstream LLMs and LVMs without requiring modifications to inference engines. Experiments demonstrate that SHiRA consistently outperforms LoRA across multiple large models: it reduces GPU memory peak usage by 16%, accelerates CPU loading speed by 5–16×, and achieves training efficiency comparable to LoRA.
为解决LoRA在视频生成中重用导致的质量下降问题,提出DART方法,通过低秩坐标变换与目标调度响应校准提高质量。
This work addresses the parameter inefficiency of conventional LoRA-based continual learning, where each task is assigned a separate adapter despite significant low-rank redundancy across tasks. The authors propose LiteLoRA, a novel approach that leverages a plug-in gating mechanism and subspace overlap analysis to selectively reuse existing adapters or instantiate new ones in a dynamic, task-aware manner. By identifying and exploiting shared low-rank subspaces among tasks, LiteLoRA achieves competitive or superior performance compared to state-of-the-art methods on standard continual learning benchmarks, while reducing the number of active adapters by 20% to 70%, thereby substantially improving parameter efficiency.
This study addresses the challenges of weight interference and retraining costs in merging multiple LoRA adapters by proposing the READ method. Built upon fixed factorization and a unidirectional coupling mechanism, READ enforces balanced canonical form constraints that enable new skills to read legacy inputs without overwriting existing outputs. This design facilitates zero-overhead skill composition during inference without requiring retraining. Experimental results demonstrate that READ surpasses baseline methods by over 20 points on benchmarks such as SuperGLUE while effectively preserving model singularity. By enabling seamless multi-skill integration within parameter-efficient fine-tuning frameworks, this work establishes a novel paradigm for adapter-based model composition.
本文针对LoRA在适应大型预训练模型时的秩利用率问题,提出了一种通过联合切空间优化的方法ISO-LoRA,以提高有效秩和下游任务表现。
Existing methods struggle to effectively merge multiple task-specific LoRA adapters into a single low-rank adapter without causing capability fragmentation or violating the low-rank structure. This work proposes a novel “Compress-then-Merge” (CtM) paradigm: it first constructs a shared r-dimensional subspace from the LoRA weights and orthogonally projects each adapter onto this subspace, then performs standard merging within the resulting r×r core coordinate space, followed by truncated SVD to strictly enforce the target rank constraint. Evaluated across multiple models and tasks, CtM significantly outperforms existing single-LoRA baselines and substantially narrows the performance gap with full-parameter merging, achieving for the first time an efficient and rank-preserving LoRA fusion.