SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing data poisoning defenses that rely on large volumes of clean validation data, which incurs high costs and is prone to overfitting in small-sample regimes. To overcome this, we propose SAGE, a framework that challenges conventional assumptions by requiring only a minimal number of mixed validation samples. By integrating a general-purpose feature extractor with non-parametric similarity-weighted prediction, SAGE accurately identifies and filters out malicious training instances. This approach substantially reduces validation overhead while significantly outperforming existing methods across multiple unlabeled attack benchmarks. Our findings demonstrate that highly effective and robust defenses against data poisoning can be achieved with remarkably few validation samples.
📝 Abstract
As machine learning increasingly relies on public, untrusted data sources, data poisoning attacks, which inject malicious examples into training data to induce misclassification of a chosen target, pose a growing threat. Existing defenses either assume zero ground-truth information about which examples are poisoned, or they assume access to a large set of examples verified to be clean. Satisfying the latter assumption incurs significant cost since reliable verification can be very resource- or labor-intensive. This cost is particularly high for clean-label attacks, where poisoned examples are visually indistinguishable from clean data. Since requiring a large set of verified examples is impractical, we propose relying on a small set of verified examples including both clean and poisoned ones, i.e., each example verified either to be clean or poisoned through inspection by a forensic expert. The challenge is then to detect poisons based on a set of verified examples that is so small that most classification models would overfit. To address this challenge, we propose Similarity-based Approach for Ground-truth-driven Exclusion (SAGE), which trains a generic feature extractor on a separate dataset and then flags poisoned training examples using a non-parametric, similarity-weighted prediction based on the verified set. On standard benchmarks against seven clean-label attack methods, we demonstrate that having access to even a handful of verified poisoned examples provides a substantial advantage. We also find that the distribution of verified clean examples across classes matters more than the number of verified examples.
Problem

Research questions and friction points this paper is trying to address.

data poisoning attacks
clean-label attacks
poison detection
verified examples
training data cleaning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Poisoning Defense
Clean-Label Attacks
Similarity-Based Detection
Non-parametric Prediction
Feature Extraction