🤖 AI Summary
This study addresses the limitations of existing exact machine unlearning methods, which rely on retraining or data sharding and suffer from low data utilization and high deletion latency. We propose Quantified Sufficient Statistics (QSS), a framework that decouples frozen patterns from mutable content to reformulate data deletion as precise arithmetic subtraction. By integrating quantized region indexing, additive statistical storage, and hybrid reconstruction strategies, QSS enables fast-path unlearning while rigorously distinguishing between two privacy guarantees: label-only deletion and complete input removal. Experimental results demonstrate that QSS achieves accuracy comparable to the SISA baseline across most tasks, while reducing deletion latency by 4× to 483× in low-class scenarios.
📝 Abstract
Exact unlearning requires a deployed predictor to match one rebuilt without the information named by a deletion request. Existing general-purpose exact methods localize retraining through disjoint shards, but every request still invalidates a model, and smaller shards reduce the data available to each constituent predictor. We introduce Quantized Sufficient Statistics (QSS), which separates a small frozen schema from mutable, sum-decomposable content. The schema learns global structure; the content stores local prediction corrections as additive statistics indexed by quantized regions. Deleting content is therefore exact subtraction rather than optimization. We distinguish two guarantees: QSS-L exactly removes a label while retaining the unlabelled input, whereas QSS-E exactly removes both input and label by learning the schema without deletable examples. A deletion takes the arithmetic fast path with probability $1-ρ$ and triggers a full rebuild with probability $ρ$; all reported expected latencies include both events. Across 15 vision, text, and tabular datasets at $ρ=0.5\%$, QSS-L is within 2 percentage points of SISA on 11 tasks and provides 4--483$\times$ lower expected deletion latency on the low-class-count tasks where a compact schema is effective. QSS-E quantifies the additional accuracy cost of removing every trace of an input.