🤖 AI Summary
Foundation model (FM)-enhanced federated learning (FL) introduces a novel backdoor attack paradigm: adversaries exploit FM vulnerabilities to inject backdoors into synthetically generated data, thereby contaminating all client models via global aggregation—diverging fundamentally from conventional backdoor attacks and evading existing defenses.
Method: We first identify anomalous hidden-layer activations as a unifying indicator across both classical and FM-enhanced backdoors. Based on this insight, we propose a dynamic activation-space constraint mechanism—requiring no access to original data and seamlessly embeddable into FL training. It comprises three components: (i) activation regularization, (ii) synthetic-data-driven constraint optimization, and (iii) latent-feature monitoring and pruning during federated aggregation.
Contribution/Results: Evaluated on multiple benchmarks, our method reduces average attack success rate to <3% while incurring <1.2% main-task accuracy degradation—significantly outperforming state-of-the-art defenses.
📝 Abstract
Federated Learning (FL) enables decentralized model training while preserving privacy. Recently, the integration of Foundation Models (FMs) into FL has enhanced performance but introduced a novel backdoor attack mechanism. Attackers can exploit FM vulnerabilities to embed backdoors into synthetic data generated by FMs. During global model fusion, these backdoors are transferred to the global model through compromised synthetic data, subsequently infecting all client models. Existing FL backdoor defenses are ineffective against this novel attack due to its fundamentally different mechanism compared to classic ones. In this work, we propose a novel data-free defense strategy that addresses both classic and novel backdoor attacks in FL. The shared attack pattern lies in the abnormal activations within the hidden feature space during model aggregation. Hence, we propose to constrain internal activations to remain within reasonable ranges, effectively mitigating attacks while preserving model functionality. The activation constraints are optimized using synthetic data alongside FL training. Extensive experiments demonstrate its effectiveness against both novel and classic backdoor attacks, outperforming existing defenses.