Score
Designs and implements training objectives, unlearning pipelines, and evaluation procedures that remove or suppress targeted adversarial behaviors, backdoor triggers, or specified internal concepts from trained models. This work uses adversarial optimization (e.g., two-stage adversarial fine-tuning), loss functions that minimize worst-case prompt influence or concept-prediction accuracy, and metrics to verify reduction of backdoor or unsafe responses.
This paper addresses security and privacy threats—such as information leakage and adversarial unlearning—in machine unlearning (MU), systematically surveying attack paradigms and defense mechanisms to fill the gap in unified threat modeling and comprehensive surveys. We propose the first four-dimensional taxonomy for MU security, categorizing attacks by target, adversary capability, operational scenario, and impact, thereby clarifying the dynamic attack-defense interplay. Our analysis integrates security assessment, privacy quantification, model inversion, and robustness evaluation, covering mainstream techniques including data removal, gradient masking, and influence function approximation. Furthermore, we introduce a verifiable threat atlas and a defense efficacy evaluation framework. The work provides theoretical foundations and practical guidelines for realizing the GDPR’s “right to be forgotten” and for delivering auditable, verifiable unlearning services in ML-as-a-Service (MLaaS) platforms.
Existing backdoor defense methods struggle to fully eliminate backdoor effects. This work addresses this challenge by, for the first time, modeling backdoor forgetting as a three-stage sequential process from the perspective of continual learning. Building upon the mechanism of catastrophic forgetting, it establishes theoretical conditions for "complete backdoor forgetting" and introduces BI-BAU, a blind inversion-based adversarial unlearning framework that operates without prior knowledge of the target class and supports multimodal contrastive learning settings. BI-BAU integrates bilevel optimization, the EM algorithm, and maximum a posteriori (MAP) estimation to embed adversarial training into the blind inversion solving procedure. Experiments demonstrate that BI-BAU effectively and thoroughly eradicates backdoors across diverse attack scenarios, exhibiting strong generalizability and superior defensive performance.
This work addresses the challenge of defending object detection models against backdoor attacks in realistic scenarios where only a poisoned model, limited clean data, and no knowledge of the attack target are available. To this end, the authors propose a detection-aware adversarial fine-tuning framework that introduces a novel soft-branch minimization mechanism to jointly handle two common backdoor behaviors—misclassification and target disappearance. The framework employs a dual-objective fine-tuning loss that selectively focuses on predictions most relevant to the backdoor for precise model repair. Compatible with both CNN and Transformer architectures, the method significantly reduces attack success rates across diverse detectors while preserving strong performance on clean samples, substantially outperforming existing backdoor defenses originally designed for classification tasks.
This work addresses the limitations of existing backdoor defense methods, which often struggle to effectively identify benign samples and remove backdoors implanted via data poisoning during training. To overcome this challenge, the authors propose HARVEY, a novel approach that instead learns an easily trainable backdoor reference model to accurately distinguish poisoned samples. HARVEY integrates a loss-difference-based sample separation strategy with an anti-backdoor learning framework, substantially improving the precision of poisoned sample identification. Extensive experiments demonstrate that HARVEY consistently outperforms state-of-the-art defenses across diverse attack types, datasets, and model architectures, achieving near-perfect backdoor removal—reducing attack success rates to nearly zero—while preserving the model’s natural accuracy with minimal degradation.
This work identifies a strong positive correlation between pretraining objective strength and backdoor persistence in vision-language models: stronger pretraining objectives—e.g., higher zero-shot transfer performance—significantly degrade the effectiveness of mainstream backdoor mitigation methods such as CleanCLIP. Using the CC3M and CC6M datasets, we systematically train multiple contrastive learning models with varying objective strengths and conduct comprehensive ablation studies—including poisoned sample removal and hyperparameter sensitivity analysis—to empirically demonstrate, for the first time, CleanCLIP’s failure under strong pretraining objectives. This finding challenges the prevailing assumption of universal efficacy among existing mitigation techniques and reveals a critical trade-off between representation quality and backdoor robustness. Our results provide both theoretical insight and empirical evidence essential for designing secure pretraining paradigms and building trustworthy multimodal AI systems.
This work investigates the security vulnerabilities of “machine unlearning”—a hazardous capability of large language models (LLMs). We demonstrate that state-of-the-art unlearning methods, such as RMU, are highly susceptible to adversarial bypasses, exhibiting weaker security guarantees than conventional safety fine-tuning. First, we systematically establish that jailbreak techniques can be adapted to recover unlearned harmful capabilities. Second, we propose two novel adaptive recovery methods: (i) targeted directional removal grounded in activation-space geometry, and (ii) few-shot fine-tuning requiring only ten irrelevant examples. Experiments show that both approaches efficiently restore the majority of unlearned harmful behaviors on RMU-edited models, exposing severe robustness deficiencies in current machine unlearning mechanisms. This study provides the first systematic adversarial evaluation framework for LLM unlearning, along with empirical evidence of its fragility, thereby advancing the development of truly attack-resilient unlearning techniques.
This work addresses the critical threat of multiple unknown backdoor attacks against large language models (LLMs), a challenge inadequately handled by existing defenses that rely on known triggers and target only single backdoors. The authors propose a novel paradigm: by deliberately injecting and then unlearning a single, controllable backdoor, they exploit its cross-backdoor generalization effect to indirectly suppress numerous unknown backdoors. For the first time, the study demonstrates and validates that backdoor unlearning exhibits generalization capabilities beyond the injected trigger. To analyze the relationship between model updates during unlearning, the authors introduce techniques such as cross-activation shift distance. Extensive experiments across three major LLM families show that unlearning just one backdoor significantly weakens diverse unknown backdoors, offering an efficient and broadly applicable new approach to enhancing LLM security.
This work addresses the threat of backdoor attacks in neural networks, which exhibit normal behavior on benign inputs but execute attacker-specified actions when triggered by a specific pattern, thereby remaining highly stealthy. To counter this, the paper introduces active path analysis—a novel approach to backdoor detection and removal—that offers both interpretability and practical utility. By identifying anomalous activation paths uniquely triggered by the backdoor pattern, the method effectively localizes and eliminates the embedded backdoor. Experimental evaluation on compromised intrusion detection models demonstrates that the proposed technique accurately identifies and successfully neutralizes backdoor triggers, confirming its effectiveness and robustness.
Current defense mechanisms struggle to detect covert harmful supervision signals embedded within seemingly benign instruction-tuning data, allowing models to be surreptitiously manipulated during fine-tuning. To address this vulnerability, this work introduces a novel threat model termed “Embedded Attack” and proposes Dual-Reference Supervised Fine-Tuning (DR-SFT), which, for the first time, integrates contrastive learning principles into supervised fine-tuning. DR-SFT combines a DPO-like pairwise preference objective with token-level regularization to enable fine-grained suppression of malicious signals. Experimental results demonstrate that DR-SFT substantially enhances model robustness against embedded harmful supervision, outperforming conventional defense strategies such as data filtering while preserving task performance.
Existing black-box attack methods struggle to effectively evaluate the robustness of multi-component NLP systems under stringent constraints—specifically, binary feedback only, no gradient access, and a query budget of ten or fewer. This work proposes a dual-agent adversarial rewriting framework: an attack agent generates semantics-preserving rewrites, while a prompt optimization agent iteratively refines the attack strategy based solely on binary feedback. The approach achieves the first effective black-box attacks under such strict conditions, revealing critical links between system architecture and vulnerability, and identifying four distinct attack patterns targeting different pipeline stages. Experiments demonstrate evasion rates of 19.95%–40.34% against four LLM-based misinformation detection systems and up to 97.02% against static retrieval systems. Furthermore, defenses informed by these attack patterns reduce evasion rates by as much as 65.18%.
This work addresses the abrupt drop in robustness during fast adversarial training, commonly attributed to catastrophic overfitting. The study offers a novel interpretation of this phenomenon through the lens of backdoor mechanisms, framing it as an unlearnable task induced by weak trigger patterns, and unifies it within a theoretical framework encompassing both backdoor attacks and unlearnable examples. To validate this perspective, the authors introduce several analytical tools—including path partitioning, feature prediction discrepancy analysis, and a universal class-discriminative trigger—and propose backdoor-inspired mitigation strategies such as vanilla fine-tuning, linear probing, weight re-initialization, and constraints suppressing weight outliers. Experimental results demonstrate that the proposed approaches not only substantiate the theoretical explanation but also significantly alleviate catastrophic overfitting and enhance model robustness across diverse adversarial attacks.