Weaponizing Ground Truth: Data Poisoning Attacks by Exploiting Boundary Misalignment Between Antivirus Software and Learning-Based Detectors

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the data poisoning vulnerability inherent in machine learning (ML) malware detectors that rely on antivirus engine labels for training. We propose Bi-Iocane, a black-box poisoning framework that, for the first time, exploits the representational discrepancies between antivirus engines and ML detectors. Specifically, it applies byte-level perturbations to modify engine-sensitive bytes, inducing label flips that corrupt the training data. Experiments across 13 antivirus engines and 8 ML detectors demonstrate that a poisoning budget of merely 0.06% achieves an average misclassification rate of 92.08%, while preserving performance on clean samples. Furthermore, existing defense mechanisms prove largely ineffective against this attack. These findings expose critical security vulnerabilities within real-world ML-based malware detection supply chains.
📝 Abstract
Machine-learning (ML)-based malware detectors are commonly trained using labels obtained from antivirus (AV) engines and aggregation services (e.g., VirusTotal). This practice assumes AV-generated labels provide reliable supervision. However, small byte-level modifications can substantially alter AV verdicts while leaving the representations perceived by downstream ML detectors largely unchanged, producing label-feature inconsistencies that can contaminate training datasets and create poisoning opportunities for ML-based malware detection. We present Bi-Iocane, a black-box poisoning framework that exploits the reliance of malware-labeling pipelines on AV-generated labels. Bi-Iocane identifies AV-sensitive bytes and modifies them to induce label changes. It rewrites such bytes in malware to obtain benign labels (evasion-oriented poisoning) and injects malware-associated byte patterns into benign software to obtain malicious labels (defamation-oriented poisoning). These poisoned samples and their lightly modified variants corrupt training data and cause selected targets to be misclassified. We evaluate Bi-Iocane with 13 AV engines simulating AV aggregation services and eight ML detectors. For 30 malware and 30 benign clean targets, Bi-Iocane combines AV-specific manipulations to generate malware-to-benign and benign-to-malware poisoned samples whose all tested AV-based labels are flipped. After these poisoned samples and variants are used for downstream training, the resulting ML models misclassify 92.08% of the original clean targets on average with only a 0.06\% poisoning budget per target. Meanwhile, the poisoned models largely preserve clean-set performance, and six evaluated poisoning defenses show only limited mitigation. VirusTotal evaluation further confirms practical defamation risk and reveals potential evasion risk in real-world AV-to-ML labeling supply chains.
Problem

Research questions and friction points this paper is trying to address.

Data Poisoning
Malware Detection
Antivirus Label Misalignment
Machine Learning Security
Adversarial Attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Poisoning
Boundary Misalignment
Black-box Attack
Malware Detection
Label Flipping
J
Jieshuai Yang
College of Cryptology and Cyber Science, Nankai University, Tianjin 300350, China
Z
Zhi Wang
College of Cryptology and Cyber Science, Nankai University, Tianjin 300350, China
Yan Jia
Yan Jia
Nankai University
IoT SecurityVulnerability DiscoverySystem SecurityNovel Attacks
Z
Zhenhua Wu
College of Cryptology and Cyber Science, Nankai University, Tianjin 300350, China
J
Jianfei Tang
College of Cryptology and Cyber Science, Nankai University, Tianjin 300350, China
C
Chenbin Su
College of Cryptology and Cyber Science, Nankai University, Tianjin 300350, China
J
Jingwei Ye
College of Cryptology and Cyber Science, Nankai University, Tianjin 300350, China
Jianwen Tian
Jianwen Tian
School of Computing and Information Systems, Singapore Management University, 80 Stamford Road, Singapore 178902
Wanpeng Li
Wanpeng Li
School of Computer Science and Informatics, University of Liverpool, Liverpool, UK