Score
Design and implement methods that fine-tune diffusion generative models with Low-Rank Adaptation (LoRA) adapters to synthesize training examples for data augmentation in few-shot or extreme-scarcity scenarios. Build and evaluate pipelines that apply LoRA-guided conditioning to control sample diversity and style, integrate the generated data into training regimes, and analyze the augmentation’s effect on downstream model performance and robustness.
Facing the challenge of inefficient navigation and filtering among over 100,000 LoRA adapters hosted on large-scale platforms, this paper proposes a submodular optimization-based adapter selection framework. It formulates adapter retrieval as a combinatorial optimization problem balancing relevance and diversity. Methodologically, we design a differentiable submodular objective function that jointly incorporates low-rank decomposition features from attention layers and semantic similarity metrics, enabling end-to-end optimization. We evaluate the approach quantitatively (retrieval accuracy, diversity score) and qualitatively (generation quality, stylistic coverage) on text-to-image generation tasks. Experiments demonstrate significant improvements across multiple domains: +12.3% Recall@10 in adapter retrieval and +28.6% LPIPS diversity gain, while preserving generation fidelity. This work establishes a novel paradigm for efficient utilization of large-scale lightweight adapter repositories.
Existing LoRA adapters require retraining when the base model changes and rely on original or high-quality synthetic data, severely limiting cross-model generalization. This paper proposes the first training-free, data-free cross-model LoRA transfer method: by enforcing subspace alignment constraints and layer-wise weight similarity filtering, it enables zero-shot mapping and reuse of LoRA parameters across distinct base models (e.g., Stable Diffusion v1.5 and SDXL). The approach comprises three core components: low-rank parameter projection, inter-layer similarity measurement, and architecture-aware alignment. Experimental results demonstrate that, without access to any original or synthetic data, the transferred adapters retain over 98% of the original LoRA’s generation performance. This significantly enhances the portability and practical utility of LoRA adapters across diverse diffusion architectures.
To address overfitting, poor generalization, and limited generation diversity in single-image fine-tuning of diffusion models, this paper proposes a timestep-dependent Low-Rank Adaptation (LoRA) framework. Our method dynamically modulates the rank constraint across timesteps and incorporates orthogonal initialization to ensure parameter independence of the adapters, effectively mitigating overfitting—particularly at high timesteps. Compared to standard LoRA and other personalization approaches, our method achieves a superior trade-off between concept fidelity and text–image alignment. Empirically, it maintains strong generalization and high-fidelity generation even under extreme data scarcity (e.g., only one input image), while remaining computationally efficient. The framework is thus well-suited for real-world deployment where both training data and computational resources are severely constrained.
To address the high power consumption, low efficiency, and memory bottlenecks in mobile LoRA fine-tuning of large diffusion models, this work proposes the first full-precision quantized training acceleration architecture specifically designed for LoRA fine-tuning. The architecture introduces a high-utilization dataflow supporting irregular tensor shapes, integrating full-precision quantized training, a customized LoRA-specific dataflow, and hardware-software co-optimization for low-rank adaptation. Experiments demonstrate 1.81× training speedup and 5.50× energy efficiency improvement over baseline implementations, with negligible degradation in image generation quality (ΔFID < 0.3). The core contribution lies in the first end-to-end integration of full-precision quantized training and LoRA-aware hardware acceleration, establishing a scalable paradigm for efficient diffusion model fine-tuning on edge devices.
Existing multi-concept customization methods for text-to-image generation suffer from attribute entanglement when fusing multiple LoRA models, necessitating separate fine-tuning to preserve concept distinguishability. This work proposes the first unified LoRA fusion framework that requires no additional fine-tuning. Its core innovation is a contrastive learning–driven weight-space alignment mechanism: by constructing positive and negative sample pairs, it optimizes the mapping of LoRA parameters to enable interference-free, high-discriminability end-to-end integration of multiple concepts within a shared diffusion model. The method jointly incorporates LoRA adaptation, contrastive loss regularization, and diffusion model fine-tuning. Experiments demonstrate substantial improvements in image fidelity and concept controllability for cross-scene compositional generation, while rigorously maintaining the semantic independence of each personalized concept.
This work addresses the lack of theoretical understanding regarding the generalization gap between Low-Rank Adaptation (LoRA) and full fine-tuning. Within a simplified linear regression framework, the study systematically compares their generalization behaviors by integrating statistical learning theory with excess risk analysis. It establishes, for the first time, a theoretical guarantee that when the discrepancy between the pre-trained model and the downstream task exhibits a low-rank structure, LoRA achieves lower excess risk than full fine-tuning in both over-parameterized and under-parameterized regimes. Empirical experiments further corroborate this finding, demonstrating a non-intuitive improvement in test accuracy under low-rank constraints and thereby revealing the intrinsic mechanism underlying LoRA’s superior generalization capability.
This work addresses the risk of training data memorization in diffusion models fine-tuned with Low-Rank Adaptation (LoRA), which poses significant copyright and privacy concerns—particularly when only LoRA weights are shared, rendering existing mitigation strategies ineffective. To tackle this challenge, the authors propose Base-Anchored Filtering (BAF), a post-processing framework that operates without access to original training data or additional retraining. BAF leverages spectral decomposition to map LoRA updates into channel space and selectively preserves generalizable features while suppressing memorized components based on their alignment with the principal subspace of the pre-trained backbone model. Notably, BAF achieves memory reduction using only the LoRA weights themselves. Extensive experiments across multiple datasets and diffusion architectures demonstrate that BAF substantially mitigates memorization risks while maintaining or even enhancing generation quality.
This work investigates whether low-rank adaptation (LoRA) can maintain strong generalization under structural constraints and proposes a more efficient, sparsity-aware fine-tuning approach. To this end, we introduce Cheap LoRA (cLA), along with its stochastic and cyclic chain variants, which enhance LoRA with sparsity. Theoretically, we establish the first information-theoretic generalization error bound for sparse LoRA and formalize cLA as a structured instance of asymmetric LoRA. Methodologically, our approach integrates sparse low-rank decomposition, randomly fixed-factor training, and cyclic parameterization. Extensive experiments across 10 pretrained models and 14 datasets demonstrate that cLA matches the performance of standard LoRA at equivalent parameter counts while reducing training time by up to 10% and peak GPU memory consumption by as much as 15%.
This work addresses the high deployment overhead and parameter interference caused by cascaded acceleration modules in existing multi-effect LoRA approaches, which often lead to concept entanglement and style degradation. To overcome these limitations, the authors propose a multi-teacher policy distillation framework that efficiently compresses up to 50 distinct visual effects and few-step generation capabilities into a single LoRA adapter. The method integrates low-rank fine-tuning, prompt-space orthogonalization, and on-policy learning through a probabilistic dual-stream routing mechanism, an asymmetric orthogonal prompt strategy, and a coarse-to-fine distillation objective. This approach significantly reduces deployment costs while achieving concept fidelity on par with or surpassing that of individual teacher models, thereby enabling effective isolation of visual effects and enhanced generalization.
This work addresses the limitation of fixed-rank constraints in parameter-efficient fine-tuning, which fail to accommodate the heterogeneous rank requirements across different layers of neural networks. The authors propose LR-LoRA, a novel approach that introduces a learnable rank mechanism within the LoRA framework, enabling differentiable and dynamic optimization of the rank for each adapter layer. This method reveals a systematic disparity in rank demands between attention and MLP layers in Transformers, thereby providing a more flexible and effective inductive bias. Experimental results demonstrate that LR-LoRA significantly outperforms existing parameter-efficient fine-tuning methods across multiple benchmarks for language understanding and commonsense reasoning, achieving state-of-the-art performance.