Score
Designs and implements lightweight adapter modules and associated insertion points for pretrained models—e.g., bottleneck, scale-and-shift (SS), post-embedding, and synergistically inserted layers—and integrates them into frozen backbones or branches to inject task-specific or modality-specific capacity. Develops and evaluates parameter-efficient tuning workflows (online/streaming updates, few-shot/few-frame learning, camera-specific and tabular adapters, cross-modal adapter fusion/integration), choosing architectures and training protocols to achieve close-to full fine-tuning performance while training far fewer parameters and analyzing fusion, temporal consistency, and bias/noise trade-offs.
Full-parameter fine-tuning of large language models (LLMs) and vision-language models (VLMs) suffers from prohibitive computational costs, overfitting, and catastrophic forgetting. Method: We propose the first structured taxonomy of parameter-efficient fine-tuning (PEFT), encompassing additive, selective, reparameterized, hybrid, and unified frameworks, and conduct the first standardized cross-modal (language/vision) and cross-task (understanding/generation) evaluation. Contribution/Results: Through theoretical analysis and multi-domain transfer experiments, we comprehensively benchmark mainstream PEFT methods—including LoRA, Adapter, and Prompt Tuning—demonstrating up to 95% GPU memory reduction and substantial computational savings while retaining over 90% of full fine-tuning performance. We further uncover fundamental trade-offs among robustness, scalability, and interpretability across PEFT paradigms, establishing a methodological foundation and empirical basis for efficient, reliable, and generalizable multimodal model adaptation.
Large language models (LLMs) face significant challenges in full-parameter fine-tuning under constrained GPU memory and computational resources, hindering efficient adaptation to downstream tasks. To address this, this work systematically surveys parameter-efficient fine-tuning (PEFT) methodologies and proposes the first unified conceptual framework—comprehensively covering theoretical foundations, algorithmic taxonomies (e.g., LoRA, Adapter, Prompt/Prefix Tuning), cross-modal extensions, and emerging trends. Distinct from fragmented surveys, our framework explicitly articulates theoretical interconnections and practical applicability boundaries across methods, unifying representative paradigms from both NLP and multimodal learning. We further release an open-source, structured knowledge graph encoding these insights. The resulting framework substantially lowers the barrier to lightweight LLM adaptation, offering researchers and practitioners a reusable, transferable technical guide. By bridging theoretical analysis with engineering pragmatism, this work accelerates the transition of PEFT from methodological exploration to scalable, production-ready deployment.
To address insufficient cross-modal fusion and excessive parameter overhead in multimodal model fine-tuning, this paper proposes the Heterogeneous Mixture-of-Experts Adapter (Heterogeneous MoE Adapter), the first to integrate heterogeneous Mixture of Experts into parameter-efficient fine-tuning (PEFT) for multimodal models. Under frozen backbone parameters, our method introduces modality-aware low-rank affine experts and a dynamic gating mechanism to enable efficient cross-modal collaboration and deep semantic fusion. Evaluated on eight vision-audio and vision-text downstream tasks, it achieves state-of-the-art performance while tuning only 5–8% of the total parameters—significantly outperforming unimodal PEFT baselines. This demonstrates the critical role of heterogeneous expert modeling in enhancing multimodal representation learning.
Existing parameter-efficient fine-tuning (PEFT) methods—such as LoRA—rely on heuristic adapter architectures, suffering from poor generalization and limited transferability across models. Method: We propose the first learnable rectangular adapter search framework grounded in Monarch matrices—the first application of Monarch structure to PEFT—supported by theoretical analysis demonstrating superior expressivity over LoRA. Our approach employs differentiable neural architecture search to automatically discover optimal lightweight adapter topologies, eliminating manual specification of rank or module shape, and integrates low-parameter adapter design with efficient fine-tuning strategies. Contribution/Results: On multi-task and multi-model benchmarks, our method significantly outperforms state-of-the-art PEFT approaches, achieving comparable or superior performance using only 5% of LoRA’s parameters. It delivers both strong cross-task/model generalization and exceptional parameter efficiency.
This work addresses the high computational cost of full fine-tuning and the limited adaptability of frozen encoders in table–image multimodal learning. To this end, the authors propose TI-Adapter, a parameter-efficient framework that freezes pretrained encoders while introducing lightweight, modality-specific adapters—comprising dedicated embedding and bottleneck layers for tables and images, respectively. The study innovatively designs distinct adapter architectures for each modality and systematically investigates the trade-offs between performance and efficiency when inserting these adapters at different network positions. Extensive experiments across 20 table–image datasets demonstrate that TI-Adapter achieves performance on par with or even superior to full fine-tuning, using only a minimal number of trainable parameters.
In adapter-based fine-tuning, input images contain confounding factors—such as pose, expression, and illumination—that entangle target attributes (e.g., identity) with spurious visual features, degrading generalization and text fidelity. To address this, we propose a counterintuitive paradigm: “introducing shortcuts to eliminate shortcuts.” During training, we incorporate removable auxiliary modules (e.g., ControlNet or LoRA) to explicitly model and isolate non-target visual factors; at inference, these modules are discarded, retaining only the lightweight adapter. Crucially, our method repurposes auxiliary modules as *disentanglement guides*—not performance enhancers—enabling clean, attribute-pure learning of the target identity. Evaluated on face and full-body identity injection tasks, our approach significantly improves generation quality, diversity, and prompt adherence. Extensive experiments demonstrate strong cross-scenario effectiveness and superior generalization, establishing the first framework to leverage auxiliary modules for explicit visual factor disentanglement in diffusion-based personalization.
This work addresses the challenges of multi-task adaptation and merging in parameter-efficient fine-tuning. We propose a modular architecture based on sparse adapters, integrating low-rank decomposition with structured weight sparsity to enable task-specific adapter training with minimal parameter updates—eliminating the need for downstream fine-tuning and supporting plug-and-play multi-task merging. Our key contribution is a streamlined, more efficient sparse training mechanism, both conceptually and empirically superior to LoRA and full-parameter fine-tuning. Experiments demonstrate the first successful joint merging of adapters across 20 diverse NLP tasks, yielding substantial gains in in-distribution performance. The approach consistently outperforms LoRA and full-model merging baselines in multi-task merging scenarios. However, cross-task generalization remains limited and warrants further investigation.
This work addresses the high computational cost of conventional fine-tuning for instance segmentation, which typically requires updating a large fraction (40–55%) of parameters in large pre-trained models. To improve parameter efficiency, the study explores parameter-efficient fine-tuning (PEFT) methods, introducing LoRA into deformable attention mechanisms for the first time and systematically evaluating the trade-offs between performance and efficiency based on the number and placement of adapters within the Transformer architecture. Experimental results demonstrate that by fine-tuning only 1–6% of the model parameters, the proposed approach matches or even surpasses the performance of full fine-tuning across four benchmark datasets. These findings validate the efficacy and feasibility of PEFT for instance segmentation and further reveal that its effectiveness is influenced by dataset complexity and model architecture.
This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.
Large language models (LLMs) face significant challenges in task adaptation under resource-constrained and closed-source API settings, where conventional parameter-efficient fine-tuning (PEFT) methods are inapplicable due to their reliance on direct model parameter access and high computational overhead. Method: This paper proposes a lightweight, parameter-free knowledge injection framework that enables task-specific adaptation without accessing the LLM’s internal parameters. Its core innovation is the “Specialized Small Model (SSM) Collaboration Paradigm,” integrating knowledge distillation from the LLM, distribution-aware task modeling, and zero-parameter coupling between the SSM and the LLM. Contribution/Results: Experiments demonstrate that our approach matches PEFT-level performance across diverse downstream tasks while reducing GPU memory consumption by over 90% and inference latency by 85%. Crucially, it operates entirely within black-box API environments—requiring no model weights, gradients, or architectural access—thus enabling seamless integration with proprietary, closed-source LLM APIs.
This work addresses the high communication and storage overhead of existing parameter-efficient fine-tuning (PEFT) methods, such as LoRA, in resource-constrained settings. The authors propose SOLAR, a model-agnostic framework that leverages subspace similarity between the base model and fine-tuning updates to reparameterize adapters as linear combinations within the subspace spanned by the base model’s singular vectors. This yields a compact adapter representation decoupled from the original PEFT architecture. SOLAR is compatible with diverse PEFT approaches and incorporates singular vector subspace modeling, controlled perturbation, and error-bound analysis to preserve performance. Experiments demonstrate that SOLAR substantially compresses adapter size across language and vision tasks on models including LLaMA, GPT, and ViT, achieving near-lossless accuracy and enabling efficient deployment in distributed and edge environments.
Existing Kronecker adapter architectures predominantly rely on fixed or heuristic designs, and the impact of their component dimensions and counts on model performance remains underexplored. This work presents the first systematic investigation into how component structure critically influences adaptation capability, introducing Configurable Decomposed Kronecker Adapter (CDKA)—a method that enables flexible structural design through parameter-budget-aware component configuration and training stability strategies. Extensive experiments across diverse natural language processing tasks demonstrate that CDKA substantially enhances adaptation performance. The study further provides a reproducible open-source implementation alongside practical deployment guidelines to facilitate broader adoption and future research.