lightweight robustness adaptation

Designs and implements small, plug-and-play modules or parameter-efficient updates that adapt pretrained models to improve robustness (e.g., to distribution shifts or noise) while training only a small subset of parameters and avoiding full end-to-end retraining. Evaluates and balances robustness gains, parameter and compute overhead, and compatibility with existing model architectures and deployment constraints.

lightweightrobustnessadaptation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Adapt Once, Thrive with Updates: Transferable Parameter-Efficient Fine-Tuning on Evolving Base Models

Jun 07, 2025
NG
Naibin Gu
🏛️ Chinese Academy of Sciences | University of Chinese Academy of Sciences | UCAS

Foundation model updates cause severe performance degradation in existing parameter-efficient fine-tuning (PEFT) modules, necessitating costly retraining. Method: This work first identifies that feed-forward networks (FFNs) encode update-sensitive knowledge, while attention mechanisms retain stable task-specific patterns. Leveraging this insight, we propose a transferable PEFT architecture that decouples task patterns from base-model knowledge. Our approach comprises attention stability analysis, FFN knowledge sensitivity modeling, parameter-decoupled structural design, and theoretical convergence guarantees. Results: Evaluated across seven foundation models and twelve datasets, our method enables zero-cost migration of legacy PEFT modules to updated models, achieving >98% average performance retention—significantly reducing operational overhead during model iteration. The core contribution is the first retraining-free PEFT transfer paradigm explicitly designed for foundation model version evolution.

PEFT modules degrade after base model updatesRe-tuning PEFT modules is computationally expensiveTrans-PEFT maintains performance without re-tuning

Robustness Feature Adapter for Efficient Adversarial Training

Aug 25, 2025
QW
Quanwei Wu
🏛️ Dongguan University of Technology | The Hong Kong University of Science and Technology (Guangzhou)

To address the dual challenges of high computational cost and robust overfitting in adversarial training (AT) for large backbone models, this paper proposes Feature-space Adapter-based Adversarial Training (FA-AT). FA-AT embeds lightweight adapter modules into intermediate feature layers, shifting adversarial perturbation injection and robust optimization from the input space to the feature space—thereby eliminating repeated, expensive gradient computations on raw inputs. Leveraging the low-rank structure of adapters, FA-AT implicitly regularizes the robust learning process, mitigating overfitting. Integrated with PGD, FA-AT achieves 35–52% training speedup on ResNet and ViT backbones, while improving robust accuracy by 2.1–4.8 percentage points on average. Moreover, it significantly enhances generalization to unseen attacks.

Eliminating robust overfitting to improve convergence qualityGeneralizing adversarial robustness to unseen attacks efficientlyReducing computational overhead in adversarial training for large models

See Further for Parameter Efficient Fine-tuning by Standing on the Shoulders of Decomposition

Jul 07, 2024
CS
Chongjie Si
🏛️ Shanghai Jiao Tong University

The theoretical mechanisms underlying parameter-efficient fine-tuning (PEFT) methods for large pre-trained models remain poorly understood, and the performance disparities among existing approaches lack principled explanations. Method: This paper establishes, for the first time, a unified theoretical framework grounded in matrix decomposition, revealing that diverse PEFT methods fundamentally perform optimization under low-rank constraints. Leveraging this insight, we propose two novel PEFT methods and a general-purpose enhancement framework—designed with theoretical rigor and architectural generality—through SVD- and LoRA-style modeling analysis, modular design, and multi-task empirical validation. Contribution/Results: Our approach significantly improves the performance of canonical PEFT methods—including LoRA and Adapter—across mainstream NLP benchmarks. This work provides the first principle-level, systematic explanation of PEFT and establishes an extensible technical pathway for future advancements.

Parameter-Efficient Fine-TuningPerformance OptimizationPre-trained Models

In large-scale pretraining, learning rate scheduling critically influences both training efficiency and model performance. This work proposes two paradigms—Fitting and Transfer. The Fitting paradigm establishes, for the first time, a scaling law for learning rate search factors, reducing hyperparameter tuning complexity from O(n³) to O(n·C_D·C_η). The Transfer paradigm extends μTransfer to Mixture-of-Experts (MoE) architectures and generalizes it across multiple hyperparameter dimensions, including depth, weight decay, and token length. Empirical results demonstrate that while μTransfer exhibits limited scalability in large-scale settings, the Fitting paradigm—grounded in the derived scaling law—offers superior scalability and practicality, providing a systematic guideline for hyperparameter tuning in industrial-scale pretraining.

hyperparameter optimizationlarge-scale pre-traininglearning rate

Meta-Learning Adaptable Foundation Models

Oct 29, 2024
JL
Jacob L. Block
🏛️ The University of Texas at Austin

Standard fine-tuning of foundation models suffers from low downstream adaptation efficiency and fails to recover the optimal adaptable parameter set. Method: We propose the first PEFT co-optimization framework that explicitly integrates meta-learning (MAML-style) into the foundation model’s retraining phase, using a LoRA-inspired low-rank adaptation structure. Contribution/Results: We theoretically prove that standard retraining is inherently suboptimal in adaptability, whereas our method strictly recovers the optimal adaptable parameters and provides a generalization error bound. Experiments on RoBERTa with the ConvAI2 dialogue continuation task demonstrate significant improvements in zero-shot and few-shot rapid adaptation performance, empirically validating the theoretical guidance.

Analyzing theoretical guarantees for parameter-efficient fine-tuning on unseen tasksDemonstrating performance improvements over standard retraining methods empiricallyDeveloping provable meta-learning for low-rank adaptation of foundation models

Latest Papers

What's happening recently
View more

Large language models (LLMs) face significant challenges in task adaptation under resource-constrained and closed-source API settings, where conventional parameter-efficient fine-tuning (PEFT) methods are inapplicable due to their reliance on direct model parameter access and high computational overhead. Method: This paper proposes a lightweight, parameter-free knowledge injection framework that enables task-specific adaptation without accessing the LLM’s internal parameters. Its core innovation is the “Specialized Small Model (SSM) Collaboration Paradigm,” integrating knowledge distillation from the LLM, distribution-aware task modeling, and zero-parameter coupling between the SSM and the LLM. Contribution/Results: Experiments demonstrate that our approach matches PEFT-level performance across diverse downstream tasks while reducing GPU memory consumption by over 90% and inference latency by 85%. Crucially, it operates entirely within black-box API environments—requiring no model weights, gradients, or architectural access—thus enabling seamless integration with proprietary, closed-source LLM APIs.

Adapts large models to tasks without accessing their parameters.Enhances performance on specific distributions using small models.Reduces resource costs for fine-tuning in constrained environments.

This work addresses the limitation of conventional adversarial training, which operates within a fixed parameter space and overlooks the impact of optimization order on model robustness. The authors propose GRAPE, a novel framework that formulates robust training as a progressive evolution in parameter space. GRAPE dynamically allocates capacity to high-stress modules guided by adversarial spectral utilization and integrates parameter-space stabilization with progressive expansion mechanisms to enable efficient training under a controlled architecture. On CIFAR-10, GRAPE achieves a PGD-20 robust accuracy of 56.94% with ResNet-18—surpassing prior methods—while reducing the number of parameters by 21.4% and incurring only a 0.9% increase in computational overhead.

Adversarial TrainingCompact ModelsOptimization Order

Existing theoretical frameworks struggle to explain why larger-scale pre-trained models substantially reduce sample complexity on downstream tasks. This work proposes a novel theoretical framework—termed “caulking”—inspired by parameter-efficient fine-tuning methods such as adapters, low-rank adaptation, and partial fine-tuning. It establishes, for the first time, a provable relationship between the scale of pre-trained models and the sample complexity of downstream tasks. By rigorously linking stronger pre-training capabilities to reduced data requirements in transfer learning, this study not only addresses a critical gap in current theoretical understanding but also provides a solid foundation for empirically observed scaling laws, demonstrating that enhanced pre-training capacity can significantly decrease the number of samples needed for effective downstream adaptation.

downstream taskspre-trained modelssample complexity

Conditional updates of neural network weights for increased out of training performance

Dec 03, 2025
JS
J. Saynisch-Wagner
🏛️ GFZ Helmholtz Centre for Geosciences

Neural networks often suffer significant performance degradation under distributional shifts—such as out-of-distribution generalization, spatiotemporal extrapolation, and cross-domain transfer—due to mismatches between training and deployment data distributions. To address this, we propose a conditional dynamic weight update framework. Its core innovation is a weight anomaly regression mechanism: sensitive weight change patterns induced by distribution shifts are identified via subset retraining; an interpretable regression predictor is then constructed to map input features to weight increments; finally, model parameters are conditionally extrapolated. The method integrates weight difference extraction, regression modeling, and extrapolation techniques, and is empirically validated on multi-source climate observation datasets. Across temporal, spatial, and cross-domain extrapolation tasks, it substantially improves prediction accuracy and robustness on out-of-distribution data, while preserving interpretability and practical applicability.

Address pattern and regime shifts in training versus application dataEnable temporal, spatial, and cross-domain extrapolation of neural networksEnhance neural network performance on out-of-distribution data

Existing parameter reparameterization methods are often confined to a single objective—either parameter-efficient fine-tuning or model compression—making it challenging to simultaneously address both demands under resource constraints. This work proposes CRISP, a unified framework that jointly achieves model compression and parameter-efficient fine-tuning within a single architecture. CRISP decomposes pre-trained weights into shared base matrices and lightweight mixture coefficients, enhanced by cross-layer base sharing and an interpolation-based gated coefficient recombination mechanism. Requiring fewer than 200 trainable parameters, CRISP outperforms existing approaches by 1% in joint compression and fine-tuning tasks, surpasses state-of-the-art methods by up to 1.5% in pure parameter-efficient fine-tuning, and achieves a consistent 4–5% improvement in overall dual-task performance.

Edge DeploymentModel CompressionNeural Network Compression

Hot Scholars

ZG

Zhiqiang Gao

Wenzhou-Kean University | Kean University
SD

Sheng Di

Argonne National Labratory, IEEE Senior Member
HPCData CompressionResilienceCloud/Grid Computing/P2P
RX

Ruiyang Xia

Xidian University
Deepfake detectionObject detectionImage steganography
ZX

Zeke Xie

Assistant Professor, The Hong Kong University of Science and Technology (Guangzhou)/ PI, xLeaF Lab
Generative AIData-centric AILarge ModelsDeep Learning Theory
LC

Longze Chen

Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Natural Language Processing