vision-language adapter fine-tuning

Designs and applies lightweight adapter-based fine-tuning methods (e.g., LoRA or adapter modules) to pre-trained vision–language models to adapt their multimodal representations to new tasks. Builds and evaluates adapter-tuned VL systems that perform outputs such as classifications, counts, categories, and localized outputs (e.g., normalized bounding boxes), and analyzes accuracy, efficiency, and transfer behavior of the adapters.

vision-languageadapterfine-tuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

$α$-LoRA: Effective Fine-Tuning via Base Model Rescaling

Oct 24, 2025
AE
Aymane El Firdoussi
🏛️ EPFL | TII

To address the insufficient generalization of pretrained models when fine-tuned on few-shot, high-dimensional binary classification tasks, this paper proposes a novel fine-tuning framework based on weight matrix reparameterization. The method couples low-rank adaptation (LoRA) with a base-model rescaling mechanism and employs random matrix theory to model the generalization behavior of high-dimensional classifiers, thereby revealing how rescaling governs spectral distribution and generalization bounds. Theoretically, the approach substantially mitigates overfitting by controlling the effective rank and condition number of the classifier’s weight matrix. Empirically, it consistently improves performance across multiple binary classification benchmarks and large language model (LLM) fine-tuning tasks, demonstrating particularly pronounced generalization gains under extreme data scarcity.

Enhancing reparameterization methods for transfer learningImproving generalization ability of fine-tuned modelsValidating effectiveness through theoretical and experimental approaches

Dynamic Adaptation of LoRA Fine-Tuning for Efficient and Task-Specific Optimization of Large Language Models

Jan 24, 2025
CW
Chihang Wang
🏛️ New York University | University of Minnesota | Tulane University | The Chinese University of Hong Kong | Stevens Institute of Technology

To address the inefficiency of fine-tuning large language models (LLMs) under resource constraints and their poor adaptability to multimodal tasks, this paper proposes Dynamic LoRA—a lightweight, input-aware fine-tuning method. Its core innovations are a layer-wise importance reallocation mechanism and a differentiable adapter architecture based on low-rank decomposition. By leveraging gradient-sensitive scoring and input-feature-distribution-guided parameter reweighting, Dynamic LoRA unifies task-specific customization with cross-task generalization. Compared to static LoRA, it achieves 88.1% accuracy and 87.3% F1 score on the GLUE benchmark while incurring only a 0.1% increase in computational overhead—significantly improving the efficiency–performance trade-off. This work establishes a novel paradigm for resource-sensitive, multimodal LLM adaptation.

Efficient Large Language ModelsMultimodal Task FlexibilityResource Optimization

MoRe Fine-Tuning with 10x Fewer Parameters

Aug 30, 2024
WT
Wenxuan Tan
🏛️ University of Wisconsin—Madison

Existing parameter-efficient fine-tuning (PEFT) methods—such as LoRA—rely on heuristic adapter architectures, suffering from poor generalization and limited transferability across models. Method: We propose the first learnable rectangular adapter search framework grounded in Monarch matrices—the first application of Monarch structure to PEFT—supported by theoretical analysis demonstrating superior expressivity over LoRA. Our approach employs differentiable neural architecture search to automatically discover optimal lightweight adapter topologies, eliminating manual specification of rank or module shape, and integrates low-parameter adapter design with efficient fine-tuning strategies. Contribution/Results: On multi-task and multi-model benchmarks, our method significantly outperforms state-of-the-art PEFT approaches, achieving comparable or superior performance using only 5% of LoRA’s parameters. It delivers both strong cross-task/model generalization and exceptional parameter efficiency.

Enhancing performance with fewer parameters in PEFT techniquesOptimizing adapter architectures for parameter-efficient fine-tuningReducing reliance on heuristics in low-rank adapters (LoRA)

This work addresses the limitations of existing Vision-Language-Action (VLA) models in cross-environment and cross-task fine-tuning, where fixed low-rank adaptation methods like LoRA fail to accommodate the high and dynamically varying intrinsic rank requirements. To overcome this, we propose LoRA-SP, the first energy-driven dynamic rank selection framework for VLA fine-tuning. LoRA-SP employs an input- and layer-adaptive rank allocation mechanism that automatically identifies critical adaptation directions based on a spectral energy threshold. By integrating SVD-style parameterization, non-negative routing scores, a shared vector bank, and energy-based pruning, our method achieves performance on par with or superior to full fine-tuning in real-world robotic multi-task manipulation—using fewer trainable parameters—and improves task success rates by up to 31.6% over standard LoRA, while significantly enhancing generalization and robustness to perturbations.

Multi-task LearningParameter-Efficient Fine-TuningRank Adaptation

Latest Papers

What's happening recently
View more

This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.

adapter placementdominant adaptation modulegradient energy

This work addresses the high computational cost of conventional fine-tuning for instance segmentation, which typically requires updating a large fraction (40–55%) of parameters in large pre-trained models. To improve parameter efficiency, the study explores parameter-efficient fine-tuning (PEFT) methods, introducing LoRA into deformable attention mechanisms for the first time and systematically evaluating the trade-offs between performance and efficiency based on the number and placement of adapters within the Transformer architecture. Experimental results demonstrate that by fine-tuning only 1–6% of the model parameters, the proposed approach matches or even surpasses the performance of full fine-tuning across four benchmark datasets. These findings validate the efficacy and feasibility of PEFT for instance segmentation and further reveal that its effectiveness is influenced by dataset complexity and model architecture.

instance segmentationlarge pretrained modelsparameter-efficient fine-tuning

This work addresses the lack of theoretical understanding regarding the generalization gap between Low-Rank Adaptation (LoRA) and full fine-tuning. Within a simplified linear regression framework, the study systematically compares their generalization behaviors by integrating statistical learning theory with excess risk analysis. It establishes, for the first time, a theoretical guarantee that when the discrepancy between the pre-trained model and the downstream task exhibits a low-rank structure, LoRA achieves lower excess risk than full fine-tuning in both over-parameterized and under-parameterized regimes. Empirical experiments further corroborate this finding, demonstrating a non-intuitive improvement in test accuracy under low-rank constraints and thereby revealing the intrinsic mechanism underlying LoRA’s superior generalization capability.

excess riskfine-tuninggeneralization

Although LoRA fine-tuning is widely adopted for large language models, the mechanisms by which it alters internal representational structures remain poorly understood. This work proposes a delta activation framework, integrated with sparse autoencoders (SAEs), to systematically isolate and analyze adapter-specific representations introduced by LoRA within the residual stream. Leveraging cosine similarity, principal angles, and centered kernel alignment (CKA), the study reveals that LoRA-induced feature dictionaries exhibit weak geometric alignment with pre-trained features, while adapter-specific SAEs achieve superior reconstruction performance. Feature density increases with both rank and network depth, yet geometric discrepancies remain stable across different ranks, suggesting that fine-tuning may generate novel representational structures not readily captured by existing interpretability tools.

feature geometryfine-tuned language modelsLoRA

This work addresses the limitation of fixed-rank constraints in parameter-efficient fine-tuning, which fail to accommodate the heterogeneous rank requirements across different layers of neural networks. The authors propose LR-LoRA, a novel approach that introduces a learnable rank mechanism within the LoRA framework, enabling differentiable and dynamic optimization of the rank for each adapter layer. This method reveals a systematic disparity in rank demands between attention and MLP layers in Transformers, thereby providing a more flexible and effective inductive bias. Experimental results demonstrate that LR-LoRA significantly outperforms existing parameter-efficient fine-tuning methods across multiple benchmarks for language understanding and commonsense reasoning, achieving state-of-the-art performance.

adapter rankinductive biaslow-rank adaptation

Hot Scholars

QZ

Qichao Zhang

中国科学院自动化研究所
人工智能 强化学习 博弈论 自适应动态规划
YZ

Yupeng Zheng

Institute of Automation, Chinese Academy of Sciences
DZ

Dongbin Zhao

Institute of Automation, Chinese Academy of Sciences
Deep Reinforcement LearningAdaptive Dynamic ProgrammingGame AISmart driving
ZX

Zebin Xing

University of Chinese Academy of Sciences
Autonomous DrivingEmbodied AI