Subtract or Replay? Exact Deletion from Language-Model Memory

📅 2026-07-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of achieving precise, targeted data removal from the persistent memory of large language models. To accommodate diverse memory representations, the authors propose a "subtract–replay" dichotomous framework: algebraically decomposable memories are updated via subtraction, while entangled representations are reconstructed through deterministic replay. This approach achieves bit-level exact unlearning for the first time in billion-scale models—including Gemma 3 (1B–12B) and the 48B Kimi Linear model—demonstrating that post-deletion model outputs exhibit no statistically significant differences from those of models never exposed to the target data, as measured by KL divergence (as low as 5.4×10⁻¹⁵), perplexity, and robustness against LiRA membership inference attacks, thereby validating both the efficacy and security of the deletion mechanism.
📝 Abstract
Exact deletion from persistent language-model memory depends on how that memory represents a record. Addressable influence can be removed by algebraic decrement; influence transformed by later writes inside shared recurrent state requires rebuilding from before the write. We test this distinction in two pretrained models against explicit record-omitted references. First, we replace Gemma 3's global-attention layers with support-vector memory. After low-rank recovery at 1B, decrement and retained-key refit agree at the next-token output to median KL $5.4\times10^{-15}$ over 31 support-token deletions, with $+2.0\%$ perplexity relative to a matched fine-tune. A masked-refit proxy is indistinguishable from the never-ingested floor under elicitation, relearning, sampling, and LiRA attacks. At 4B and 12B, certificate ordering persists but utility cost rises to $11.2\%$ and $44.3\%$. Second, in a 48B Kimi Linear hybrid, additive writes admit a fixed decrement and diagonal decay a corrected one, whereas the delta rule makes $12$--$49\%$ of a record's contribution suffix-dependent. Checkpointed rewind-and-replay deletes real clinical records at contexts up to 18,842 tokens, matching never-ingested logits and all recurrent states bit for bit within a deterministic MLX implementation; replaying a correction provides exact amendment. Exact deletion is therefore a property of memory representation: subtract addressable records and replay entangled writes.
Problem

Research questions and friction points this paper is trying to address.

exact deletion
language-model memory
data removal
memory representation
machine unlearning
Innovation

Methods, ideas, or system contributions that make the work stand out.

exact deletion
support-vector memory
rewind-and-replay
addressable memory
language model unlearning
🔎 Similar Papers
2024-10-03arXiv.orgCitations: 0