🤖 AI Summary
This study addresses the challenge of machine unlearning in large language models, wherein eliminating the influence of target data often compromises retained capabilities within complex systems such as retrieval-augmented generation, long-term memory, and multi-agent frameworks. To tackle this, the work proposes a five-layer architecture linking removal requests to supporting claims alongside a seven-stage lifecycle framework, establishes a six-dimensional evidence evaluation taxonomy, and systematically reviews existing unlearning algorithms across these complex settings. The findings reveal that target construction and recovery testing critically determine unlearning efficacy. Furthermore, the authors demonstrate that current evidence remains insufficient to substantiate influence removal across external states, underscoring the necessity of dependency tracking to verify whether the influence of targeted data re-emerges over time.
📝 Abstract
Machine unlearning aims to remove target influence while preserving other capabilities. This survey compares methods, benchmarks, and evidence across large language models and systems using retrieval, memory, tools, and interacting agents. A five-layer framework connects removal requests, system boundaries, target locations, interventions, and supported claims. A seven-stage lifecycle and six evidence dimensions guide comparison. The review shows that target construction, retained data, and recovery tests affect reported outcomes. Evidence from model evaluations remains insufficient to establish removal across external state and subsequent updates, motivating evaluation that tracks dependencies and tests whether target influence returns.