Score
Designs and implements lightweight adapter modules and tuning protocols that attach to pretrained visual foundation models so the base encoder can remain frozen; builds training recipes and parameter-efficient mechanisms that minimize additional trainable weights while enabling adaptation to new downstream vision tasks or modalities and preserving model performance.
To address the computational and memory bottlenecks inherent in full-parameter fine-tuning of billion- or trillion-parameter vision foundation models, this work systematically investigates parameter-efficient fine-tuning (PEFT) methods for vision. We formally define vision PEFT for the first time and propose a unified taxonomy comprising three categories: additive (e.g., LoRA, Adapter), selective (e.g., BitFit), and unified (e.g., VPT, Prompt Tuning). Through comprehensive evaluation across diverse pretraining paradigms and cross-task generalization benchmarks, we survey state-of-the-art approaches, standard datasets, and critical open challenges. Our study establishes the most complete knowledge framework for vision PEFT to date, accompanied by an open-source repository covering over 100 works. This resource provides both a theoretical foundation and practical guidance for efficient vision transfer learning.
Full-parameter fine-tuning of large language models (LLMs) and vision-language models (VLMs) suffers from prohibitive computational costs, overfitting, and catastrophic forgetting. Method: We propose the first structured taxonomy of parameter-efficient fine-tuning (PEFT), encompassing additive, selective, reparameterized, hybrid, and unified frameworks, and conduct the first standardized cross-modal (language/vision) and cross-task (understanding/generation) evaluation. Contribution/Results: Through theoretical analysis and multi-domain transfer experiments, we comprehensively benchmark mainstream PEFT methods—including LoRA, Adapter, and Prompt Tuning—demonstrating up to 95% GPU memory reduction and substantial computational savings while retaining over 90% of full fine-tuning performance. We further uncover fundamental trade-offs among robustness, scalability, and interpretability across PEFT paradigms, establishing a methodological foundation and empirical basis for efficient, reliable, and generalizable multimodal model adaptation.
Existing parameter-efficient fine-tuning (PEFT) methods—such as LoRA—rely on heuristic adapter architectures, suffering from poor generalization and limited transferability across models. Method: We propose the first learnable rectangular adapter search framework grounded in Monarch matrices—the first application of Monarch structure to PEFT—supported by theoretical analysis demonstrating superior expressivity over LoRA. Our approach employs differentiable neural architecture search to automatically discover optimal lightweight adapter topologies, eliminating manual specification of rank or module shape, and integrates low-parameter adapter design with efficient fine-tuning strategies. Contribution/Results: On multi-task and multi-model benchmarks, our method significantly outperforms state-of-the-art PEFT approaches, achieving comparable or superior performance using only 5% of LoRA’s parameters. It delivers both strong cross-task/model generalization and exceptional parameter efficiency.
Large-scale pretrained models face high computational overhead and structural instability during multi-task adaptation. Method: This paper proposes a composable fine-tuning framework that integrates graph-structured task priors with modular adapters. It constructs a task-relation graph to model inter-task dependencies, leveraging this structured prior to guide low-rank adapter parameter allocation and dynamic routing. The framework incorporates plug-and-play adapter design, relation-matrix regularization, and temperature- and gating-based control mechanisms to mitigate path conflicts and redundant computation. Contributions/Results: Experiments demonstrate significant improvements in task prediction accuracy and adapter assignment precision. The method exhibits strong robustness under hyperparameter, environmental, and data perturbations, achieving both high performance and parameter efficiency. It establishes a new paradigm for multi-task adaptation—characterized by interpretability, reusability, and structural stability—without compromising scalability or practicality.
To address the prohibitive computational and memory overhead of full-parameter fine-tuning for large language models, this paper proposes Quantum-inspired Adapter (QAdapter), the first method to deeply integrate Hamming-weight-preserving structure, orthogonality constraints, and matrix composition mechanisms into a low-dimensional adapter module enabling multi-order feature coupling. Crucially, QAdapter rigorously maintains parameter orthogonality while substantially enhancing representational capacity and generalization. Experiments demonstrate that QAdapter achieves 99.2% of full fine-tuning performance on GLUE and VTAB benchmarks, using only 1/44 the parameters of LoRA. Moreover, with merely 1/25 the parameters of OFT/BOFT, it attains 98% of their relative performance. This work establishes a novel paradigm for efficient large-model adaptation by unifying structural invariance, geometric constraints, and compositional expressivity in lightweight parameter-efficient tuning.
The theoretical mechanisms underlying parameter-efficient fine-tuning (PEFT) methods for large pre-trained models remain poorly understood, and the performance disparities among existing approaches lack principled explanations. Method: This paper establishes, for the first time, a unified theoretical framework grounded in matrix decomposition, revealing that diverse PEFT methods fundamentally perform optimization under low-rank constraints. Leveraging this insight, we propose two novel PEFT methods and a general-purpose enhancement framework—designed with theoretical rigor and architectural generality—through SVD- and LoRA-style modeling analysis, modular design, and multi-task empirical validation. Contribution/Results: Our approach significantly improves the performance of canonical PEFT methods—including LoRA and Adapter—across mainstream NLP benchmarks. This work provides the first principle-level, systematic explanation of PEFT and establishes an extensible technical pathway for future advancements.
To address the lack of input awareness and task-specific modeling capability in existing parameter-efficient fine-tuning (PEFT) adapters, this paper proposes a dynamic input-conditioned Transformer architecture. The core innovation is the input-Conditioned Network (iCoN), which generates instance-specific, channel-wise dynamic convolutional kernels for fine-grained, input-adaptive feature modulation. Our method fine-tunes only 1.6%–2.8% of the backbone parameters, yet achieves full fine-tuning performance on depth estimation and semantic segmentation, and significantly outperforms full fine-tuning on image classification and instance segmentation. It consistently surpasses mainstream PEFT approaches—including LoRA and Adapter—across diverse downstream tasks. By enabling input-aware, task-adaptive representation learning with minimal parameter overhead, the proposed method substantially enhances the generalization capability and expressive power of PEFT across heterogeneous vision tasks.
This work addresses the challenge of efficiently constructing and managing massive numbers of persistent personalized models atop trillion-parameter foundation models. It proposes leveraging parameter-efficient fine-tuning (PEFT) as a lightweight and reliable personalization substrate, combining a shared large model with small, trainable adapters to encode user preferences, skills, and memory. The authors introduce MinT, an infrastructure that integrates adapter identity management, version control, provenance tracking, evaluation, and serving mechanisms, and define three scaling dimensions: Scale Up, Scale Down, and Scale Out. Experimental results demonstrate that, even under strong shared priors, compact adapters can stably capture personalized behaviors, offering a viable pathway toward large-scale deployment of millions of persistent personal models.
This work addresses the limitation of existing image compression methods that neglect the joint optimization of statistical and semantic information in entropy models when adapting pretrained codecs, thereby constraining the effectiveness of parameter-efficient fine-tuning. To overcome this, the authors propose S2-CoT, a structure–semantics co-tuning framework that systematically analyzes and coordinates adapter type and placement. Specifically, they introduce a Structure-Fidelity Adapter (SFA) for the codec and a Semantic Context Adapter (SCA) for the entropy model, enabling dual-adapter joint optimization through parameter-efficient fine-tuning, spatial–frequency feature fusion, and channel-wise context modeling. Evaluated on four mainstream codecs, S2-CoT achieves performance comparable to full fine-tuning using only a minimal number of trainable parameters, significantly enhancing compression efficiency for machine vision tasks and establishing new state-of-the-art results.
This work addresses the high communication and storage overhead of existing parameter-efficient fine-tuning (PEFT) methods, such as LoRA, in resource-constrained settings. The authors propose SOLAR, a model-agnostic framework that leverages subspace similarity between the base model and fine-tuning updates to reparameterize adapters as linear combinations within the subspace spanned by the base model’s singular vectors. This yields a compact adapter representation decoupled from the original PEFT architecture. SOLAR is compatible with diverse PEFT approaches and incorporates singular vector subspace modeling, controlled perturbation, and error-bound analysis to preserve performance. Experiments demonstrate that SOLAR substantially compresses adapter size across language and vision tasks on models including LLaMA, GPT, and ViT, achieving near-lossless accuracy and enabling efficient deployment in distributed and edge environments.
This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.
This study addresses the challenges of parameter inefficiency and knowledge fragmentation in cross-domain adaptation for visual foundation models by proposing a Self-Routed Tensor Adapter. The method integrates a shared Tucker core with learnable domain matrices, augmented by a progressive depth-weighted routing strategy that enables sample-level adaptive mixing without external gating networks. Experimental evaluations across five benchmarks demonstrate that this framework matches or exceeds the accuracy of Mixture-of-Experts baselines while utilizing less than 30% of their parameters. Consequently, this approach effectively facilitates efficient and generalizable multi-domain visual representation learning, yielding significant improvements in parameter efficiency compared to existing methods.