🤖 AI Summary
This work addresses the challenge that existing clean-label backdoor attacks struggle to simultaneously achieve low pre-unlearning attack success rates and strong post-unlearning activation under realistic machine unlearning scenarios. To this end, the authors propose a dormancy-activation backdoor framework based on dual-generator learning, which establishes sample-specific persistent latent associations coupled with a removable suppression mechanism to precisely trigger the backdoor after target samples are unlearned. This approach is the first to jointly optimize these two objectives under the constraints of clean labels and realistic unlearning. By integrating bilevel optimization, unlearning simulation, and label-consistent disguised samples, the method significantly outperforms baseline approaches on CIFAR-10 and ImageNet-10, achieving lower pre-unlearning attack success rates while enabling stronger post-unlearning attack efficacy across multiple unlearning algorithms.
📝 Abstract
Existing backdoor attacks often become effective immediately after backdoor implantation and may therefore be exposed before exploitation. Machine unlearning activated dormant backdoors mitigate such behavioral exposure by remaining inactive after training and becoming effective only after selected training records are unlearned. However, existing methods struggle to simultaneously achieve a low pre-unlearning attack success rate and strong post-unlearning activation under clean-label constraints and realistic unlearning requests. Achieving this transition requires jointly establishing a persistent latent association and a removable suppressive influence. To address this challenge, we propose a clean-label unlearning-activated backdoor framework based on dual-generator learning and formulate it as a bilevel optimization problem: By simulating latent backdoor establishment and machine unlearning, the framework alternately learns sample-specific triggers that establish a latent trigger-to-target association and label-consistent camouflage samples that provide removable suppression. Once a small subset of camouflage samples is unlearned, the suppression is lifted and the dormant backdoor is activated. Experiments on CIFAR-10 and ImageNet-10 show that our method maintains lower pre-unlearning attack success rates while achieving stronger post-unlearning activation across multiple unlearning algorithms than representative backdoor baselines. These results demonstrate that reliable dormancy-to-activation transitions can be achieved by coordinating a persistent latent association with removable suppression under clean-label and realistic deletion constraints.