Score
Designs, implements, or evaluates methods for adapting a pretrained model backbone to new tasks while updating only a small fraction of its parameters to avoid full fine‑tuning. This includes techniques such as selective unfreezing, adapter modules, low‑rank or sparse updates, and other parameter‑efficient mechanisms that reduce trainable parameter count while preserving backbone weights.
Large language models (LLMs) face significant challenges in full-parameter fine-tuning under constrained GPU memory and computational resources, hindering efficient adaptation to downstream tasks. To address this, this work systematically surveys parameter-efficient fine-tuning (PEFT) methodologies and proposes the first unified conceptual framework—comprehensively covering theoretical foundations, algorithmic taxonomies (e.g., LoRA, Adapter, Prompt/Prefix Tuning), cross-modal extensions, and emerging trends. Distinct from fragmented surveys, our framework explicitly articulates theoretical interconnections and practical applicability boundaries across methods, unifying representative paradigms from both NLP and multimodal learning. We further release an open-source, structured knowledge graph encoding these insights. The resulting framework substantially lowers the barrier to lightweight LLM adaptation, offering researchers and practitioners a reusable, transferable technical guide. By bridging theoretical analysis with engineering pragmatism, this work accelerates the transition of PEFT from methodological exploration to scalable, production-ready deployment.
To address the computational and memory bottlenecks inherent in full-parameter fine-tuning of billion- or trillion-parameter vision foundation models, this work systematically investigates parameter-efficient fine-tuning (PEFT) methods for vision. We formally define vision PEFT for the first time and propose a unified taxonomy comprising three categories: additive (e.g., LoRA, Adapter), selective (e.g., BitFit), and unified (e.g., VPT, Prompt Tuning). Through comprehensive evaluation across diverse pretraining paradigms and cross-task generalization benchmarks, we survey state-of-the-art approaches, standard datasets, and critical open challenges. Our study establishes the most complete knowledge framework for vision PEFT to date, accompanied by an open-source repository covering over 100 works. This resource provides both a theoretical foundation and practical guidance for efficient vision transfer learning.
Large-scale pretrained models face high computational overhead and structural instability during multi-task adaptation. Method: This paper proposes a composable fine-tuning framework that integrates graph-structured task priors with modular adapters. It constructs a task-relation graph to model inter-task dependencies, leveraging this structured prior to guide low-rank adapter parameter allocation and dynamic routing. The framework incorporates plug-and-play adapter design, relation-matrix regularization, and temperature- and gating-based control mechanisms to mitigate path conflicts and redundant computation. Contributions/Results: Experiments demonstrate significant improvements in task prediction accuracy and adapter assignment precision. The method exhibits strong robustness under hyperparameter, environmental, and data perturbations, achieving both high performance and parameter efficiency. It establishes a new paradigm for multi-task adaptation—characterized by interpretability, reusability, and structural stability—without compromising scalability or practicality.
To address structural distortion and topological inconsistency in high-dimensional parameter spaces (e.g., 4D tensors) induced by low-rank approximation in parameter-efficient fine-tuning, this paper proposes a structure-preserving low-rank core space modeling method. Unlike conventional low-rank adapters (e.g., LoRA), which are restricted to linear weight matrices, our approach explicitly models and preserves the intrinsic topological structure of the original high-dimensional parameter space—achieving compact and accurate reconstruction of N-dimensional parameter updates via high-order tensor decomposition. Evaluated across CV, NLP, and multimodal benchmarks, the method yields an average accuracy improvement of 1.8% under identical parameter budgets, while reducing structural distortion by 37%, significantly outperforming existing baselines.
Large language models (LLMs) face significant challenges in task adaptation under resource-constrained and closed-source API settings, where conventional parameter-efficient fine-tuning (PEFT) methods are inapplicable due to their reliance on direct model parameter access and high computational overhead. Method: This paper proposes a lightweight, parameter-free knowledge injection framework that enables task-specific adaptation without accessing the LLM’s internal parameters. Its core innovation is the “Specialized Small Model (SSM) Collaboration Paradigm,” integrating knowledge distillation from the LLM, distribution-aware task modeling, and zero-parameter coupling between the SSM and the LLM. Contribution/Results: Experiments demonstrate that our approach matches PEFT-level performance across diverse downstream tasks while reducing GPU memory consumption by over 90% and inference latency by 85%. Crucially, it operates entirely within black-box API environments—requiring no model weights, gradients, or architectural access—thus enabling seamless integration with proprietary, closed-source LLM APIs.
Traditional model ensembling relies on averaging numerous fine-tuned models, incurring high computational cost and low efficiency. This paper proposes Model Stock: an efficient ensemble method requiring only two fine-tuned models. Its key insight is that fine-tuned weights closer to the layer-wise weight center exhibit superior in-distribution (ID) and out-of-distribution (OOD) generalization. Leveraging this, Model Stock introduces a layer-wise dual-model weighted averaging strategy designed to approximate the center of the weight space. Built upon the CLIP architecture, it jointly optimizes layer-wise center approximation and OOD robustness. On standard ID/OOD benchmarks, Model Stock consistently outperforms state-of-the-art methods—including Model Soup—achieving higher ID accuracy and stronger OOD robustness, while introducing negligible inference overhead.
The theoretical mechanisms underlying parameter-efficient fine-tuning (PEFT) methods for large pre-trained models remain poorly understood, and the performance disparities among existing approaches lack principled explanations. Method: This paper establishes, for the first time, a unified theoretical framework grounded in matrix decomposition, revealing that diverse PEFT methods fundamentally perform optimization under low-rank constraints. Leveraging this insight, we propose two novel PEFT methods and a general-purpose enhancement framework—designed with theoretical rigor and architectural generality—through SVD- and LoRA-style modeling analysis, modular design, and multi-task empirical validation. Contribution/Results: Our approach significantly improves the performance of canonical PEFT methods—including LoRA and Adapter—across mainstream NLP benchmarks. This work provides the first principle-level, systematic explanation of PEFT and establishes an extensible technical pathway for future advancements.
This work addresses catastrophic forgetting in continual learning with pre-trained models when access to previous task data is prohibited. The authors propose a structured low-rank adaptation method grounded in geometric redundancy of pre-trained weights. By analyzing the intrinsic geometric structure of the pre-trained weight space, they identify a protected subspace for parameter updates and formulate the update as \( \Delta W = BAQ^\top \), where frozen matrices \( B \) and \( Q \) project trainable low-rank matrix \( A \) exclusively onto redundant directions. This approach is the first to leverage geometric redundancy to explicitly locate plasticity regions, enabling a controllable trade-off between plasticity and stability without requiring data replay. Experimental results demonstrate that the method effectively suppresses functional drift and significantly improves retention of performance on prior tasks, even under worst-case scenarios.
This paper addresses machine unlearning—the efficient removal of a model’s dependence on specific training samples while preserving performance on the remaining data. We propose a constraint-optimization-based feasible update framework. Our core innovation introduces a parameter masking mechanism to select an updateable subspace, jointly incorporating gradient noise modeling and directional constraints on parameter updates to yield locally feasible solutions satisfying both unlearning objectives and utility preservation. The method operates as a plug-and-play module, enhancing the robustness and accuracy of diverse first-order approximate unlearning algorithms. Experiments on image classification tasks demonstrate that our approach significantly improves unlearning accuracy (average gain of 12.3%) while incurring negligible utility loss—less than 0.5% drop in test accuracy on retained data—validating its effectiveness and practicality.
To address the insufficient generalization of pretrained models when fine-tuned on few-shot, high-dimensional binary classification tasks, this paper proposes a novel fine-tuning framework based on weight matrix reparameterization. The method couples low-rank adaptation (LoRA) with a base-model rescaling mechanism and employs random matrix theory to model the generalization behavior of high-dimensional classifiers, thereby revealing how rescaling governs spectral distribution and generalization bounds. Theoretically, the approach substantially mitigates overfitting by controlling the effective rank and condition number of the classifier’s weight matrix. Empirically, it consistently improves performance across multiple binary classification benchmarks and large language model (LLM) fine-tuning tasks, demonstrating particularly pronounced generalization gains under extreme data scarcity.
Existing parameter-efficient fine-tuning methods struggle to recover the multi-scale local geometric information lost during downsampling in 3D point cloud backbone networks, thereby limiting dense prediction performance. This work proposes a local-aware parameter-efficient fine-tuning framework that, while keeping the backbone frozen, introduces for the first time a hierarchical local feature reconstruction mechanism. This mechanism leverages a multi-resolution local feature pyramid, a local-global semantic fusion module, and a dynamic multi-scale prompt generator to restore fine-grained geometric structures, coupled with a lightweight upsampling segmentation head for efficient adaptation. The approach achieves state-of-the-art performance with only 2.71% (for classification) and 7.69% (for dense prediction) trainable parameters, and scales effectively on PointGPT-L with merely 0.36% additional parameters.
This work addresses the challenge of targeted machine unlearning in large language models (LLMs). Unlike conventional fine-tuning or data-deletion approaches, we propose an efficient unlearning method grounded in the model editing paradigm. We are the first to systematically evaluate and adapt causal mediation–based editing algorithms—including ROME, IKE, and WISE—for machine unlearning tasks. Crucially, we reformulate the editing objective to emphasize precise knowledge localization and controllable, localized parameter modification—driven by gradients or activations—enabling accurate removal of targeted information. On multiple standard unlearning benchmarks, our method substantially reduces residual memory rates while limiting downstream task performance degradation to under 3%; in certain settings, it outperforms state-of-the-art unlearning baselines. Our core contributions are: (i) establishing model editing as a novel, high-fidelity paradigm for targeted unlearning; and (ii) introducing principled design criteria and technical pathways for unlearning-aware editing objectives.