Score
Designs and implements lightweight adapter modules that apply feature-wise affine transformations (learned per-channel scale and bias) to a model's intermediate activations, often conditioned on external control signals. These adapters are built to be initialized to identity so they preserve the base model prior and enable parameter-efficient fine-tuning and multi-modal control with minimal paired data.
To address the computational and memory bottlenecks inherent in full-parameter fine-tuning of billion- or trillion-parameter vision foundation models, this work systematically investigates parameter-efficient fine-tuning (PEFT) methods for vision. We formally define vision PEFT for the first time and propose a unified taxonomy comprising three categories: additive (e.g., LoRA, Adapter), selective (e.g., BitFit), and unified (e.g., VPT, Prompt Tuning). Through comprehensive evaluation across diverse pretraining paradigms and cross-task generalization benchmarks, we survey state-of-the-art approaches, standard datasets, and critical open challenges. Our study establishes the most complete knowledge framework for vision PEFT to date, accompanied by an open-source repository covering over 100 works. This resource provides both a theoretical foundation and practical guidance for efficient vision transfer learning.
Existing parameter-efficient fine-tuning (PEFT) methods—such as LoRA—rely on heuristic adapter architectures, suffering from poor generalization and limited transferability across models. Method: We propose the first learnable rectangular adapter search framework grounded in Monarch matrices—the first application of Monarch structure to PEFT—supported by theoretical analysis demonstrating superior expressivity over LoRA. Our approach employs differentiable neural architecture search to automatically discover optimal lightweight adapter topologies, eliminating manual specification of rank or module shape, and integrates low-parameter adapter design with efficient fine-tuning strategies. Contribution/Results: On multi-task and multi-model benchmarks, our method significantly outperforms state-of-the-art PEFT approaches, achieving comparable or superior performance using only 5% of LoRA’s parameters. It delivers both strong cross-task/model generalization and exceptional parameter efficiency.
Existing parameter-efficient fine-tuning (PEFT) methods for Mamba architectures over-rely on adapting the State Space Model (SSM) module, neglecting the critical role of the projector in transfer learning. Method: This work first identifies the projector as the dominant component for cross-task adaptation and proposes ProDiaL—a novel PEFT method that introduces only a learnable diagonal linear transformation applied centrally to the pretrained projector, enabling targeted, non-weight-updating adaptation while freezing all SSM parameters. Contribution/Results: ProDiaL decouples projector optimization from SSM learning, reducing trainable parameters by >99% (<1% of total). On both vision and language Mamba models, it achieves performance on par with full fine-tuning at minimal computational cost, demonstrating strong generalization. As the first projector-centric PEFT paradigm for Mamba, ProDiaL challenges the prevailing SSM-centric design philosophy and establishes a new direction for efficient Mamba adaptation.
Parameter-efficient fine-tuning (PEFT) methods, particularly LoRA, underperform for large operator models in scientific machine learning due to theoretical limitations in Fourier layers. Method: This work introduces PEFT to scientific ML for the first time and proposes F-Adapter—a frequency-adaptive adapter architecture grounded in the spectral sparsity of physical systems. It allocates parameters dynamically across the frequency domain: high capacity at low frequencies and low capacity at high frequencies. Coupled with a spectral-complexity-driven module-width optimization strategy, F-Adapter jointly enhances parameter efficiency and approximation accuracy. Contribution/Results: On multiple 3D Navier–Stokes benchmarks, F-Adapter significantly outperforms LoRA and other state-of-the-art PEFT methods, achieving new SOTA performance while improving generalization and spectral fidelity.
This work addresses the limitations of task-specific adapters in class-incremental learning, which suffer from restricted knowledge transfer, high retrieval overhead, and catastrophic forgetting induced by parameter fusion. To overcome these challenges, the authors propose a dynamic adapter fusion mechanism grounded in PAC-Bayes theory. Operating under a frozen pre-trained model, the method dynamically computes fusion coefficients via Taylor expansion and incorporates a global knowledge-aware robust initialization strategy to effectively integrate task-specific, historical global, and initial parameters. This approach balances stability and plasticity, preserving learned knowledge while enabling adaptation to new tasks. Extensive experiments demonstrate that the proposed method significantly outperforms existing approaches across multiple class-incremental learning benchmarks, achieving state-of-the-art performance.
To address the lack of input awareness and task-specific modeling capability in existing parameter-efficient fine-tuning (PEFT) adapters, this paper proposes a dynamic input-conditioned Transformer architecture. The core innovation is the input-Conditioned Network (iCoN), which generates instance-specific, channel-wise dynamic convolutional kernels for fine-grained, input-adaptive feature modulation. Our method fine-tunes only 1.6%–2.8% of the backbone parameters, yet achieves full fine-tuning performance on depth estimation and semantic segmentation, and significantly outperforms full fine-tuning on image classification and instance segmentation. It consistently surpasses mainstream PEFT approaches—including LoRA and Adapter—across diverse downstream tasks. By enabling input-aware, task-adaptive representation learning with minimal parameter overhead, the proposed method substantially enhances the generalization capability and expressive power of PEFT across heterogeneous vision tasks.
研究通过使用共享循环核心和残差适配器库学习灵活的运动基元,解决了如何学习神经科学理论中提出的系统问题。
This work addresses the challenge of efficiently adapting and aligning large language models to multiple tasks without modifying their pretrained weights. The authors propose LARA, a lightweight adaptation method that freezes the backbone model and injects low-rank correction signals into the residual stream. LARA introduces token-level dynamic routing within the residual stream for the first time, enabling concurrent hosting and on-demand composition of multiple behavioral modules. It further incorporates a tunable interpolation coefficient γ to enable smooth control over behavior blending. With only 33 MB of additional overhead, LARA supports the simultaneous deployment of seven distinct behaviors on a 1.5B-parameter model, achieving performance comparable to LoRA on code fine-tuning and DPO tasks while significantly enhancing deployment flexibility and resource efficiency.
This work addresses the high computational cost of full fine-tuning and the limited adaptability of frozen encoders in table–image multimodal learning. To this end, the authors propose TI-Adapter, a parameter-efficient framework that freezes pretrained encoders while introducing lightweight, modality-specific adapters—comprising dedicated embedding and bottleneck layers for tables and images, respectively. The study innovatively designs distinct adapter architectures for each modality and systematically investigates the trade-offs between performance and efficiency when inserting these adapters at different network positions. Extensive experiments across 20 table–image datasets demonstrate that TI-Adapter achieves performance on par with or even superior to full fine-tuning, using only a minimal number of trainable parameters.
This study addresses the performance degradation in large language model fine-tuning caused by hardware noise and precision limitations in analog compute-in-memory systems, where full retraining remains prohibitively expensive. To this end, we propose a parameter-efficient adaptation method based on Low-Rank Adaptation (LoRA). By integrating input reshaping with update accumulation techniques, our approach achieves optimizer-agnostic robust adaptation under fixed pre-trained weights, effectively overcoming errors in forward and backward matrix-vector multiplications as well as the challenges of physical weight updates under limited conductance states. Experimental results demonstrate that the proposed method significantly improves the fine-tuning performance of Llama-series models in noisy environments, maintaining superior efficacy even under stringent conditions restricted to merely 20 conductance states.
This work addresses the high computational and memory costs associated with transfer learning in speech and audio foundation models by introducing, for the first time, the Mamba state space model into a parameter-efficient transfer learning (PETL) framework. The authors propose a lightweight method based on low-rank bottleneck adapters, embedding a shared-parameter Mamba module within the adapter architecture to substantially reduce the number of trainable parameters while enhancing audio feature modeling capacity. Experimental results demonstrate that the proposed approach achieves competitive or superior performance compared to existing PETL methods across four audio classification benchmarks and speech recognition tasks spanning five languages—even under stricter parameter budgets.