ADS-C: Antidistillation Sampling for Classification

📅 2026-07-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the vulnerability of proprietary classifiers to model stealing via knowledge distillation from their soft outputs. To counter this threat, the authors propose an anti-distillation sampling method that injects input-dependent, gradient-guided perturbations into the teacher model’s output distribution. Crucially, this approach preserves the teacher’s Top-1 accuracy exactly—incurred zero utility loss—while significantly degrading the distillability of its predictions. The core innovation lies in integrating a closed-form confidence-bound budgeted perturbation, temperature scaling analysis, and gradient-oriented distributional perturbation, thereby achieving the first utility-preserving defense against distillation-based theft. Experimental results demonstrate substantial drops in student model accuracy—by 17.4, 29.6, and 13.3 percentage points on CIFAR-100, CIFAR-10, and Tiny-ImageNet, respectively—without any degradation in the teacher model’s performance.
📝 Abstract
Knowledge distillation enables an adversary to replicate a proprietary classifier by querying its prediction interface and training a surrogate on the returned probability vectors. Antidistillation sampling, proposed for large language models, counters this threat with an input-dependent, gradient-directed perturbation of the served distribution; its transfer to classification has not been studied. Adapting the defense to classification, we show its behavior is governed by the distribution of the teacher's per-input confidence margins. Because well-trained classifiers are severely overconfident, the direct transfer exhibits an inert window: below a closed-form-predictable threshold, it affects neither attacker nor defender; beyond it, the defense undergoes a phase transition and degrades the teacher faster than the attacker's student. Temperature softening rescales the transition in closed form, and every temperature configuration lies on the same unfavorable trade-off curve. Our method, ADS-C, composes the perturbation under a closed-form, per-input margin budget that provably preserves every served top-1 prediction, so the defended teacher's accuracy equals the undefended teacher's identically. Under this guarantee the distilled student still loses 17.4 percentage points on CIFAR-100, 29.6 on CIFAR-10, and 13.3 on Tiny-ImageNet; matching this degradation with the unmodified defense costs 27.5, 32.9, and 22.2 points of teacher accuracy. Because served labels are unchanged, a hard-label attacker gains nothing, while the defended soft output trains a student up to 29.7 points below that floor: the incentive to distill served probabilities is not merely removed but reversed. To our knowledge, ADS-C is the first antidistillation defense for classification whose utility cost is exactly zero.
Problem

Research questions and friction points this paper is trying to address.

knowledge distillation
model stealing
antidistillation
classification
adversarial defense
Innovation

Methods, ideas, or system contributions that make the work stand out.

antidistillation
knowledge distillation defense
confidence margin
zero utility loss
classification robustness
🔎 Similar Papers
No similar papers found.