Score
Designs and implements training protocols and model variants that fix (freeze) specified neural network parameters or subcomponents (e.g., a backbone) while updating only the remaining weights; this includes building and training lightweight adapter or MLP modules, scheduling freeze/unfreeze stages, applying alignment or transfer losses, and fine-tuning the unfrozen parameters. It also covers evaluating the impact of freezing on learned representations and decoders and transferring or benchmarking adapters across different backbones.
Large language models (LLMs) face significant challenges in full-parameter fine-tuning under constrained GPU memory and computational resources, hindering efficient adaptation to downstream tasks. To address this, this work systematically surveys parameter-efficient fine-tuning (PEFT) methodologies and proposes the first unified conceptual framework—comprehensively covering theoretical foundations, algorithmic taxonomies (e.g., LoRA, Adapter, Prompt/Prefix Tuning), cross-modal extensions, and emerging trends. Distinct from fragmented surveys, our framework explicitly articulates theoretical interconnections and practical applicability boundaries across methods, unifying representative paradigms from both NLP and multimodal learning. We further release an open-source, structured knowledge graph encoding these insights. The resulting framework substantially lowers the barrier to lightweight LLM adaptation, offering researchers and practitioners a reusable, transferable technical guide. By bridging theoretical analysis with engineering pragmatism, this work accelerates the transition of PEFT from methodological exploration to scalable, production-ready deployment.
Existing adapter-based fine-tuning methods, though parameter-efficient, suffer from substantial memory and computational overhead, prolonged training time, and imbalanced adapter contributions. This work is the first to empirically reveal the heterogeneous contribution of adapters within Transformer architectures. To address these issues, we propose a selective freezing mechanism: (i) dynamically assessing adapter importance via gradient sensitivity and task-specific contribution; (ii) implementing a staged, progressive freezing strategy; and (iii) incorporating implicit regularization to smooth the loss landscape and improve generalization. Our method maintains or even improves downstream task performance while significantly reducing memory consumption (−42.85%), FLOPs (−34.59%), and training time (−11.82%). The approach achieves superior efficiency without compromising robustness or accuracy, offering a principled and practical solution for resource-constrained adapter tuning.
To address the inefficiency, deployment overhead, and poor generalization of large language models (LLMs) in multi-task settings—including multilingual support, question answering (QA), and structured output generation—this paper proposes a multi-encoder–frozen-decoder architecture: the entire decoder is frozen, while only lightweight, task-specific encoders and adapter modules are fine-tuned. This work presents the first systematic empirical validation of decoder freezing under realistic multi-task and multilingual configurations, overcoming the limitation of prior frozen-parameter approaches—which were largely confined to single-task or representation-learning scenarios. Built upon the AlexaTM model and the PEFT paradigm, our method substantially reduces training cost and mitigates catastrophic forgetting. It maintains state-of-the-art (SOTA) performance on natural language generation (NLG) tasks, and—remarkably—outperforms full-parameter fine-tuning baselines on QA and structured generation tasks, achieving an unprecedented balance between deployment efficiency and cross-task generalization.
This work addresses the degradation in training throughput and model accuracy caused by suboptimal parameter freezing strategies in pipeline-parallel training. It presents the first unified formulation that jointly models pipeline scheduling and precision constraints: computational dependencies are represented via a directed acyclic graph, and an optimization problem is formulated as a linear program with explicit accuracy constraints. This framework dynamically determines the optimal proportion of frozen parameters to minimize per-batch execution time. The resulting approach enables adaptive, accuracy-aware freezing policies that achieve up to a 40% improvement in training throughput on the LLaMA-8B model while preserving model accuracy, and it generalizes effectively across diverse pipeline-parallel configurations.
Traditional layer freezing still requires forward propagation through frozen layers, limiting computational efficiency gains; while feature map caching holds promise, it faces two overlooked challenges: augmentation invalidation and substantial memory overhead. This paper proposes the first systematic framework to address these issues: (1) a similarity-aware channel-level data augmentation strategy to mitigate distribution shift in cached features, and (2) a lossy progressive compression scheme that significantly reduces storage cost without compromising accuracy. Experiments across diverse models (ResNet, ViT) and benchmarks (ImageNet, CIFAR) demonstrate that our approach reduces training FLOPs by 32–47%, decreases GPU memory consumption by 58–73%, and incurs only marginal accuracy degradation (0.1–0.3%). These results validate the method’s efficiency, robustness, and scalability.
The functional mechanisms underlying layer-wise operations in Transformer models remain poorly understood, particularly regarding the necessity and interchangeability of layer ordering. Method: This work proposes an empirical analysis framework based on freezing large language models (LLMs) and systematically conducts three types of architectural interventions: layer ablation, layer reordering, and parallel layer execution. Contribution/Results: We discover that middle layers exhibit strong functional uniformity and order invariance—enabling safe skipping, arbitrary reordering, or concurrent execution—thereby challenging the conventional assumption of strict layer-order dependency. Across diverse downstream tasks, skipping or parallelizing middle layers reduces inference latency by up to 30% on average, with accuracy degradation under 1%. This study is the first to empirically characterize the functional heterogeneity spectrum across Transformer layers, providing an interpretable, evidence-based foundation for model lightweighting, architectural compression, and novel variant design.
This work addresses the optimization instability and lack of theoretical capacity guidance associated with inserting adapters into frozen vision Transformer backbones during transfer learning. The authors propose the Zero-initialized Residual Low-rank Adapter, which introduces a low-rank bottleneck structure in each Transformer block, with the up-projection layer initialized to zero to ensure that fine-tuning starts identically to the pretrained model, thereby preventing early representation drift. For the first time, the adapter rank is theoretically modeled as a capacity budget tied to the feature shift of downstream tasks, revealing an “elbow”-shaped accuracy gain as rank increases. Experiments across nine datasets and three backbone scales show that the method improves top-1 accuracy by 14.9% on average over training only the classification head, using just 0.92% of the parameters required for full fine-tuning, and outperforms full fine-tuning in 10 out of 15 dataset-backbone combinations.
This work addresses the issue of silent, irreversible training stagnation in low-precision optimization, where weight updates are truncated when rounding errors fall below half a unit in the last place (ULP). The study presents and validates, for the first time, a deterministic condition for weight freezing in low-precision training: by leveraging the high-precision (fp32) training trajectory and the mantissa length of the target floating-point format (e.g., fp8 or bf16), the exact iteration at which each weight coordinate freezes can be predicted a priori—without executing actual low-precision training. Using a mantissa-truncation simulator combined with stochastic rounding, the method accurately forecasts freezing points on models such as GPT and GPT-2 with an error of at most four steps, revealing that weight freezing is the primary cause of validation loss plateaus. Stochastic rounding effectively mitigates this phenomenon, and the approach generalizes across diverse models, tasks, and floating-point formats.
This study addresses the challenge of catastrophic forgetting in sequential fine-tuning of pretrained language encoders, a phenomenon exacerbated when using full-parameter updates, while the mechanisms underlying parameter-efficient methods remain poorly understood. Through controlled experiments, we systematically evaluate the sequential learning performance of Low-Rank Adaptation (LoRA) on BERT-base and RoBERTa-base, complemented by representation probing and fine-grained ablation studies to elucidate its anti-forgetting properties. We provide the first empirical evidence that standard LoRA substantially mitigates task-level forgetting, achieving an average forgetting rate of only 0.6% ± 1.4%, markedly lower than full fine-tuning (19.9% ± 4.8%) and Elastic Weight Consolidation (15.5% ± 1.4%). This effectiveness primarily stems from freezing the backbone network, thereby preserving a stable shared feature structure across tasks.
This work investigates the proportion of parameters in neural networks that are truly necessary for encoding task-specific information and proposes a method that trains only extremely low-rank LoRA adapters while keeping the backbone network entirely frozen and randomly initialized. The approach is validated across diverse architectures and tasks, revealing that task-relevant information resides in an exceptionally low-dimensional subspace. This finding implies that randomly initialized backbones are interchangeable and need only be distributed as random seeds. By linking the saturation rank of LoRA to the intrinsic dimensionality of tasks, the method recovers 96%–100% of full fine-tuning performance using merely 0.5%–40% trainable parameters across nine benchmarks, substantially reducing storage and memory overhead.