Rethinking Backdoor Repair Evaluation: Distinguishing Aggregate Clean Utility from Benign Performance Preservation

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对后门修复评估,通过区分总体清洁效用与良性性能保留,提出使用最差类保留损失和尾部保留损失来补充整体清洁准确率,以解决局部性能退化被掩盖的问题。
📝 Abstract
Backdoor repair aims to suppress malicious behavior in compromised models while preserving benign task performance. Existing studies typically evaluate these objectives using Attack Success Rate (ASR) and Overall Clean Accuracy, but aggregate clean accuracy can obscure substantial degradation concentrated in a small portion of the label space. We revisit benign-performance evaluation from a preservation perspective by distinguishing aggregate clean utility from the preservation of previously available class-wise performance. We define class-wise preservation loss by comparing clean performance before and after repair and show that aggregation can hide localized degradation through localized-loss dilution and cross-class compensation. To complement Overall Clean Accuracy, we characterize localized preservation loss using Worst-Class Preservation Loss and Tail Preservation Loss. We conduct a systematic empirical study across representative backdoor attacks, repair methods, datasets, attack targets, and model architectures, with additional validation under clean-label attacks. Results show that effective attack suppression and favorable aggregate clean performance do not necessarily imply uniform preservation of previously available benign performance across classes. Substantial localized preservation losses can remain, and their severity and class-wise structure vary across repair conditions. These findings motivate preservation-oriented class-wise evaluation alongside ASR and Overall Clean Accuracy.
Problem

Research questions and friction points this paper is trying to address.

backdoor repair
benign performance preservation
aggregate clean accuracy
class-wise preservation loss
localized degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

class-wise preservation loss
worst-class preservation loss
tail preservation loss
🔎 Similar Papers
No similar papers found.
B
Baogang Song
School of Artificial Intelligence, Wuhan University of Technology
C
Changtian Song
School of Artificial Intelligence, Wuhan University of Technology
J
Jian Chen
School of Artificial Intelligence, Wuhan University of Technology
F
Fan He
School of Artificial Intelligence, Wuhan University of Technology
Junwei Zhou
Junwei Zhou
Dartmouth College
Computer Vision
Jianwen Xiang
Jianwen Xiang
Wuhan University of Technology
Dependable ComputingSoftware EngineeringFormal MethodsKnolwedge Management
Dongdong Zhao
Dongdong Zhao
Wuhan University of Technology
Biometrics SecurityPrivacy-preserving Deep LearningArtificial Intelligence Security